A conversation with Sachin Bhandari on data integrity when AI enters the GxP record
The ALCOA+ principles have carried data integrity in GxP for decades, and they are about to be tested by a technology their authors never imagined. Data integrity remains among the most frequently cited categories in FDA drug GMP warning letters, named in roughly 61 percent of those issued in 2021 by one European Pharmaceutical Review analysis [1], so it is the ground every inspection still turns on.
As AI begins to draft deviation narratives, summarise batch records, and shape quality decisions, those same ALCOA principles are being pushed in ways that strain the assumptions behind them.
We sat down with Sachin Bhandari, a Quality IT and validation leader, member of ISPE with more than 25 years in pharma and a track record of taking AI-supported quality systems through health-authority inspection, to ask what it really takes to keep ALCOA+ standing when the author of a record might be a model.
Key takeaways
- The principles hold, but the evidence gets harder. AI does not break ALCOA+. It increases the amount of evidence required to demonstrate compliance.
- Attribution becomes a chain. Who made the record is now a person, a model version, a prompt, and a dataset. The accountable human is whoever accepted the output.
- The original is no longer a single artefact. Reconstructing an AI-generated record may require the output, prompt, model version, timestamp, and source data.
- Accuracy becomes risk-based. Probabilistic outputs should be assessed against defined tolerances and their intended use.
- Human oversight must be meaningful. Review only works as a control when the reviewer has the time, information, competence, and authority to challenge the output.
Why is AI testing assumptions that held for decades?
ALCOA+ has been the backbone of data integrity in GxP for decades. What has made it so durable, and why is AI now testing some of the assumptions behind it?
Sachin: It lasted because it asks a human question, not a technology one: can I trust this record, and can I rebuild what happened?
That question does not much care what sits underneath it, which is exactly why it carried us all the way through paper, through spreadsheets, and through validated systems without ever really needing to change. The technology kept moving and the question stayed still.
Where AI starts to test it is in a handful of assumptions we never quite said out loud. ALCOA+ quietly assumed that a record gets created once, by a person you can name, at a moment you can point to, and that it then sits there unchanged.
Every one of those assumptions wobbles with AI.
The author might be a model rather than a person. The original is no longer a single thing you can hold; it is a bundle of the prompt, the version of the model, and the data behind it. And because the output is probabilistic, the same question asked tomorrow might come back a little differently.
So I would say the principles are as sound as they ever were. What has changed is how much evidence you have to put on the table to satisfy them.
“AI doesn’t weaken ALCOA+. It asks us to show a great deal more to prove the same thing.” — Sachin Bhandari
Throughout your career, you have witnessed several major shifts, from paperless validation and cloud adoption to the move from CSV to CSA. How does the industry’s current AI moment compare?
Sachin: Every big shift in my career moved trust to a new place, and AI moves it further than any of them.
When we went paperless, trust moved from the wet signature on the page to the audit trail sitting behind it. When we went to the cloud, it moved again, this time onto a shared line with a vendor we now depended on. And the move from CSV to CSA shifted it once more, away from the sheer weight of documentation and towards thinking honestly about risk.
AI takes that further because it moves trust into the reasoning itself.
With everything that came before, the system still did what we told it to, and our job was simply to prove that it had. With AI, the system reaches a conclusion of its own, and now we have to show that the conclusion holds up.
The good news, and I think it is genuinely good news, is that the discipline we built during the CSA years is exactly what this moment asks for: intended use, risk-based thinking, and testing the things that actually matter.
So a team that took CSA seriously is a long way from starting cold.
The part I will be honest about is that this is the first shift I have lived through where the system can quietly change its own behaviour without anybody raising a change request, and that is genuinely new ground for all of us.
“This is the first shift I’ve lived through where the system can change its own behaviour without anybody raising a change request.” — Sachin Bhandari
What does it take to translate ALCOA+ for AI?
When you try to apply the ALCOA principles to AI-driven processes, which one becomes the most difficult to defend, and why?
Sachin: Attributable is the hardest to defend, with Original close behind, and Accuracy, the one everyone expects to struggle with, turns out to be the most manageable.
Attributable gets difficult because the honest answer to who made this record is no longer a single name. It is a person, plus a particular version of a model, plus the prompt that was used, plus the data the model learned from.
Attribution turns into a chain rather than a signature.
Original is hard for a related reason. The source is no longer one clean artefact, so you actually have to decide, deliberately and in advance, what your original is going to be.
If you forced me to pick just one, I would say Attributable, because accountability is the thing regulators care about most deeply, and AI is precisely the technology that smudges it.
Some argue that ALCOA+ needs a complete rethink for AI, while others say it simply needs better interpretation. Where do you sit in that debate, and why?
Sachin: I sit firmly on the side of better interpretation, as long as it is genuine interpretation and not just a new coat of paint.
The principles are really about trust and reconstruction, and neither of those needs to change simply because the author happens to be a model rather than a person.
A complete rethink would throw away decades of hard-won understanding between industry and regulators, and worse, it would send the signal that the old discipline no longer applies, which is not true.
What we actually need is new evidence sitting underneath the same principles, plus one honest addition on top. That addition is decision integrity layered above data integrity. It adds a floor rather than replacing the house.
We have done this kind of translation before, by the way. When we went paperless, Contemporaneous did not need rewriting. We simply agreed that the timestamp in the audit trail was the contemporaneous record, and we moved on.
AI is asking us to make the very same kind of move.
When it comes to “translating” ALCOA+ for AI, what does that actually look like in practice?
Sachin: In practice, it comes down to asking one simple question of every principle: what would I actually show an inspector now?
For Attributable, that means capturing the prompt, the model and its version, and the person who accepted the output, so the audit trail reaches all the way to the model.
For Original and Contemporaneous, it means timestamping and pinning the version at the moment the output was generated, and keeping the prompt and the model’s state right alongside the result.
For Accurate, it means moving the question from “Is this correct?” to “Is this correct within a tolerance we have defined, and who is checking the cases that really matter?”
None of this is exotic, and that is rather the point.
It shows up in the same places good practice has always lived: the URS, the intended-use statement, the audit-trail design, and the monitoring plan. The skeleton is the one CSA already gave us, and we are simply hanging model-aware evidence on it.
To make it concrete, instead of storing only the line “root cause: equipment fault”, the record now also holds the prompt, the model version, the timestamp, and the name of the analyst who accepted it.
The conclusion on the page has not changed at all, but you can stand behind it two years later when somebody asks.
Figure 1. ALCOA+ translated for AI across four areas: training data, model artefact, version traceability, and electronic signatures.
What does Attribution look like when an AI system contributes to a decision?
Sachin: It becomes a chain of accountability, and the word I would underline is contributes, because the AI never owns the decision. A person always does.
At the very least, you want to capture who ran it, what the input or prompt was, which model and which version produced the output, and who then reviewed it and accepted it into the GxP process.
The accountable human is the one who accepted the output, not the one who built the model, and I would make that explicit, because that is the name that belongs on the record.
The way this goes wrong is an AI output drifting into a batch record with no thread leading back to the model, the version, or the prompt that produced it.
A record like that cannot really be attributed to anyone, and it will not hold up when an inspector pulls on it.
I find the simplest way to think about it is to treat the AI like a junior analyst’s draft. The junior can absolutely write the thing, but a named, qualified person reads it, signs it, and owns it.
“Think of the AI as a junior analyst’s draft. It can write it, but a named, qualified person signs it and owns it.” — Sachin Bhandari
When does human review meaningfully reduce risk?
One of the most common controls proposed for AI is human review. In your experience, when does human oversight meaningfully reduce risk, and when is it simply a stamp we put on?
Sachin: Human review only reduces risk when the reviewer can actually catch the error, has what they would need to catch it, and has a real route to say no.
Take any one of those away and it quietly becomes a stamp.
It turns into a stamp the moment someone is signing off more than they could possibly read, or when they cannot see the reasoning behind the AI’s output, or when they have no genuine authority to overrule it.
A human in the loop on an organisational chart is not, by itself, a control.
To make the review mean something, you give the reviewer the inputs and some sense of how confident the model was, you tell them clearly what they are looking for, and you make it easy to record a disagreement.
A review process where nobody ever overrules anything is not reassuring to me. It is a warning sign.
And I will be honest: a good deal of what passes for human oversight today is closer to the stamp than to the real thing, and inspectors are getting noticeably better at telling the two apart.
Picture a model screening four hundred records a day and one reviewer approving all four hundred inside an hour. There is simply no way that person read them.
Real oversight looks like the model handing the dozen genuinely uncertain cases to a human who has the time to look at them properly.
“A review process where nobody ever overrules anything isn’t a comfort. It’s a warning sign.” — Sachin Bhandari
How should Accuracy be defined for probabilistic AI outputs?
Accuracy has always been a cornerstone of data integrity. How should organisations think about accuracy when AI outputs are probabilistic rather than deterministic?
Sachin: I would define accuracy as fitness for the intended use within a tolerance you have defined, rather than as a single right answer, and honestly, we already work this way in other corners of quality.
You set the acceptance criteria before you deploy, and you tie them to the risk of what the model is doing.
A model that drafts a summary for a human to review can live with far more uncertainty than one feeding an automated release decision, and the criteria should reflect that difference.
Then you test against a known, representative dataset, you state plainly the performance you are prepared to accept, and you keep watching to see that it holds.
The shift in mindset is to stop asking whether the thing is always right and start asking whether it is right often enough, in the ways that matter, and what happens on the occasions when it is wrong.
We never demanded that an analytical method be perfect. We set limits for accuracy and precision, and we validated against them.
A probabilistic model is really the same idea, just applied to a new kind of system.
What do Original and Contemporaneous mean for AI?
The terms Original and Contemporaneous were developed around human-created records. What happens when an AI generates a deviation narrative, a batch-record summary, or a risk assessment? Is the prompt part of the record? The model version? Both?
Sachin: Both, and that is really the whole answer.
The record is the output together with enough context to rebuild how it came to exist, which at the very least means the prompt and the model version.
Contemporaneous means recorded at the time of the event, and for AI the event is the generation itself, so the timestamp, the prompt, and the state of the model all belong to that moment.
Original means the first capture of the data, and if you keep only the tidy final output, you have kept the conclusion and quietly thrown away the evidence that backs it.
The test I give teams is an easy one to remember.
If you could not regenerate the output, or at the very least explain exactly how it came about, six months later and in front of an inspector, then you have not really captured the original.
Take an AI that writes a deviation narrative. Keep only the paragraph and a year from now you cannot say what facts it was working from or even which model wrote it.
Keep the prompt, the version, and the timestamp alongside it, and you can reproduce the whole thing and stand behind it without hesitation.
ALCOA+ still holds, but the evidence has changed
AI does not require the industry to abandon the principles that have governed pharmaceutical data integrity for decades. It requires organisations to interpret those principles against a more complex chain of authorship, evidence, and accountability.
The record is no longer only the final paragraph, classification, or recommendation produced by the system. It may also include the prompt, the model version, the source data, the timestamp, the acceptance criteria, and the qualified person who reviewed and accepted the result.
That is what translating ALCOA+ for AI ultimately means: keeping the principles intact while expanding the evidence required to prove them.
The next challenge is what happens beyond the record itself.
Once an AI output begins influencing a GxP decision, organisations must be able to demonstrate not only that the record is trustworthy, but that the model remained in its validated state, the training data was governed, vendor changes were controlled, and the resulting decision was sound. This is the territory the draft EU GMP Annex 22 on artificial intelligencehas begun to map, alongside the revised Annex 11 and Chapter 4.
That is where decision integrity, model drift, shadow AI, and inspection readiness enter the picture.
If your organisation is working through what AI means for its validated systems, BGO Software’s regulatory and compliance advisory team works at exactly this intersection. Get in touch to continue the conversation.
About Sachin Bhandari
Sachin Bhandari is a Quality IT and digitalisation leader with more than 25 years in pharma, working at the intersection of CSV, CSA, and the validation of AI and ML systems in GxP environments.
He has led enterprise eQMS, paperless validation, and cloud quality programmes at global scale, and has taken AI-supported quality systems through health-authority inspection in live operations.
He is the founder of TrustBridge Compliance and the author of Validating AI in GxP: A Practitioner’s Guide. More at trustbridge-compliance.com.
Reference notes
Reference notes (for editorial verification, current as of June 2026)
[1] Data integrity among the most frequently cited categories in FDA drug GMP warning letters — European Pharmaceutical Review: “FDA warning letters highlight data integrity issues” (61 percent of 2021 letters) and RAPS: “Experts offer advice on avoiding common warning letter citations”. See also the FDA “Data Integrity and Compliance With Drug CGMP” guidance.
[2] EU GMP draft Annex 22 (AI), Annex 11 and Chapter 4 — ECA Academy / GMP-Compliance: drafts released for comment and the European Commission consultation hub.
[3] What Annex 22 spells for AI in GMP — European Pharmaceutical Review: “What Annex 22 spells for AI in GMP manufacturing”.
[4] EMA 2026 GMP revisions (Chapter 4, Annex 11, Annex 22) — Epista: “Preparing for EMA’s 2026 GMP revisions”.
[5] ALCOA++ and GxP data integrity — IntuitionLabs: “ALCOA+ Principles: A Guide to GxP Data Integrity”. For the foundational regulator definitions, see the MHRA “Guidance on GxP data integrity” and WHO Technical Report Series No. 996, Annex 5, “Guidance on Good Data and Record Management Practices.”



