Which version of your model is in production, and who approved it?
Your data team knows which version is the winning one. What an audit asks is something else: who decided to put it into production, what controls were in place, and exactly which files were running on the day of the incident.
What MLflow Already Knows, and What It Doesn't
MLflow is the de facto standard for logging models. It tracks which run produced each version, what data was used for training, and what metrics were obtained. That's enough for the data team.
In terms of governance, three things are missing. Evidence-based controls. Pre-production approval, because right now anyone with write access can change the production alias without leaving a record. And proof that today’s files are the same ones that were in production on the day of the audit.
Six controls, reproducible results
Verifiable AI MLOps Governance reads the company's MLflow—without ever writing to it—and evaluates each version using six checks without any language models:
Input diagram.
The model should specify what data it expects to receive.
Serialization.
Do not save it in formats that can execute code when loaded.
Requirements.
Libraries that are version-locked and have no known security advisories.
Data lineage.
Which datasets were used to train each version?
Description and License.
What is the model, and under what license is it used?
Promotion.
Make sure that whatever is in production has been approved.
In the MLflow demo, a properly governed model passes 6 out of 6. Another model, in production without approval, fails 0 out of 6 and triggers an alert. The difference isn’t based on opinion—it’s determined by a rule. With the same inputs, rules, and evaluation sources, the result is reproducible.
Before production, someone says yes
The deployment approval is signed by a responsible person, along with their attestation. If a release reaches production without it, an alert is triggered in Monitoring; the next round of checks verifies the approval and updates the alert’s status. This is the human oversight required by Article 14, applied where the risk originates.
The Model V-Seal: What Exactly Was That Version?
Stamping a version records the fingerprint of each file in the model, its datasets, and its results, with cryptographic proof. If a metric or a file changes, a new stamp is required.
The company only releases an encrypted manifest containing fingerprints and counts. Neither the model, nor the data, nor the value of any metric leaves its infrastructure. Anyone can verify the seal on the public page without an account.
Why It Matters Now
The requirements set forth in the European AI Regulation for a high-risk system (data, documentation, record-keeping, oversight, and accuracy: Articles 10, 11, 12, 14, and 15) stem from the model’s lifecycle. That is where the evidence either exists or does not exist.
To be honest: synchronization is triggered manually, so a change to an alias is detected on the next run, not immediately. And with MLflow on Databricks, fingerprint calculation hasn't been tested yet; it's handled by the pipeline script.
Every version is tested. Every deployment is approved and verifiable.
A model, its versions, and a review of each one.
Recorded on the V-PROOF Portal connected to a demo MLflow instance.
Watch the demo