A QA Framework for ML Models: How I Built ModelContract
Published: 2026-10-03 · Author: AI Release · @ai_release1
⚡ The Gist in 5 Seconds - The author developed a framework for automated testing of pre-trained ML models: it builds a ModelContract from the model's metadata and formal API. - The experience is based on the Olist model (late delivery risk assessment) with preserved metrics: ROC AUC 0.8101, PR AUC 0.3354, Brier 0.0633. - The main limitation: PASS/FAIL criteria must be explicitly defined in the metadata or in the model itself; the dataset is not a source of requirements. ### 🔍 What Was Discovered While trying to test her own ML model, the author ran into a problem typical of Data Science: the documentation exists only in the developer's head, and the dataset readily shows actual values but doesn't answer the question of "how things should be." A minimum in the data doesn't mean an acceptable boundary, and the absence of a feature combination is not the same as its being invalid. Formal criteria are needed for verification. The solution was found in a ready-made Olist model: alongside it was documentation, including feature descriptions, a dataset, preserved metrics, and library versions. This made it possible to move from "checking documentation" to implementing a framework. The key idea is to separate the human-readable README from the machine-readable model_metadata.json. The Olist metadata specifies the target "is_late," test metrics, and versions of scikit-learn 1.9.0, pandas 3.0.5, and numpy 2.4.6. Additionally, the framework reads information directly from the loaded model: feature_names_in_, classes_, and the presence of predict and predict_proba. Thus ModelContract was born — a set of requirements assembled from two sources that remembers the origin of each item. ### 💡 Why It Matters The main value of this approach is that the checks are not tied to a specific model. The framework contains only the checking mechanism, while all requirements come from the contract of whatever project is loaded into it. This makes it possible to automate the testing of any pre-trained ML models without thinking about the domain. And it's precisely the combination of metadata and the model's API that makes it possible to find discrepancies: for example, if the metadata declares 34 features but feature_names_in_ returns 33, that immediately becomes a concrete FAIL. The plan for the future is to add information about who made a change and when to the report, so you can go straight to the culprit with questions. An example of such a FAIL: "The test dataset lacks the feature distance_km," requirement from spec v2, item 4.1, change made by @ivanov_dev, PR #142 dated 15.09.2026. ### 🧩 Context Six months ago, the author had already created an agent for checking technical specifications. The new task — ML testing — arose from a practical need: her own project turned out to have no documentation, making it an ideal candidate for the experiment. The absence of formal criteria in her own project led to choosing Olist as the guinea pig, and then to building the framework. It currently implements five checks and is ready to be used on other models.
⚡ The Gist in 5 Seconds - The author developed a framework for automated testing of pre-trained ML models: it builds a ModelContract from the model's metadata and formal API.
- The experience is based on the Olist model (late delivery risk assessment) with preserved metrics: ROC AUC 0.8101, PR AUC 0.3354, Brier 0.0633.
- The main limitation: PASS/FAIL criteria must be explicitly defined in the metadata or in the model itself; the dataset is not a source of requirements.
🔍 What Was Discovered While trying to test her own ML model, the author ran into a problem typical of Data Science: the documentation exists only in the developer's head, and the dataset readily shows actual values but doesn't answer the question of "how things should be." A minimum in the data doesn't mean an acceptable boundary, and the absence of a feature combination is not the same as its being invalid.
Formal criteria are needed for verification.
The solution was found in a ready-made Olist model: alongside it was documentation, including feature descriptions, a dataset, preserved metrics, and library versions.
This made it possible to move from "checking documentation" to implementing a framework.
The key idea is to separate the human-readable README from the machine-readable model_metadata.json.