Get In Touch

Beyond Data Integrity: Model Drift, Shadow AI, and Defending AI Decisions in GxP

Updated - 09 Sep 2026 17 min read
bgosoftware logo
BGO Software The Digital Health Lab

A conversation with Sachin Bhandari on model drift, shadow AI, training data, vendor oversight, and inspection readiness

Applying ALCOA+ to an AI-generated record is only one part of the challenge.

The greater risk may sit beyond the record itself: in the model’s changing behaviour, or model drift, the data it learned from, the commercial platform that updates it, the employee using an unapproved chatbot, or the decision that a reviewer accepts without genuinely understanding how it was reached.

As AI begins influencing deviation assessments, batch-record reviews, risk classifications, and quality decisions, organisations must be able to demonstrate more than data integrity. They must show that the model was still operating within its validated state and that the decision it helped shape was sound, controlled, and defensible. Call that second obligation decision integrity.

We spoke with Sachin Bhandari, a Quality IT and validation leader with more than 25 years in pharma and experience taking AI-supported quality systems through health-authority inspection, about the risks organisations are still missing and what AI inspection readiness will require.

Key takeaways

  • Decision integrity is the new layer. Proving a record is accurate is no longer enough. Organisations must be able to defend the decision the AI helped shape.
  • Model drift is a data integrity risk. A model behaving differently from the day it was validated has quietly left its validated state.
  • Shadow AI is already entering GxP records. Unapproved tools can introduce content with no attribution, model version, or audit trail.
  • Training data is a GxP input. If the model supports GxP decisions, the data it learned from requires lineage, governance, and impact assessment.
  • Vendor models create silent-change risk. Embedded AI may be updated without the regulated organisation being informed.
  • Inspection readiness requires the full decision story. The evidence package must connect the input, model, version, monitoring, human review, and final accountable decision.

Why does model drift create a data integrity risk?

Traditional validation assumes that once a system is validated, its behaviour stays stable until a controlled change is introduced. That assumption breaks down with AI. How should organisations think about model drift from a data-integrity and inspection-readiness perspective?

Sachin: I would treat drift as a data integrity risk, not just a number on a performance dashboard.

If a model is behaving differently today from the way it behaved on the day you validated it, then your validated state has quietly expired, and the uncomfortable part is that nobody signed for that change. It happened on its own.

So the practical thing is to decide, before you ever go live, what still counts as the validated model.

That means being clear about the kind of data you expect to see coming in, the performance you are willing to hold the model to, and the points at which something should raise its hand and alert you.

Then you watch it for the life of the system, because validation now runs across the whole lifecycle instead of finishing neatly at go-live.

An inspector is not going to stop at asking whether you validated the model.

The sharper question, and the one I would prepare for, is how you know it still does today what it did on the day you validated it, and what would trigger you to go back and look again.

The trap I see most often is treating retraining as quiet housekeeping.

It is not.

Retraining is a change. It deserves the same control as any other change, and you want a record of what the model looked like before and after.

“If the model behaves differently today than it did the day you validated it, your validated state has quietly expired, and nobody signed for the change.” — Sachin Bhandari

Are organisations focusing on the right AI risks?

In your view, are organisations focusing on the right risks when introducing AI into GxP environments? What are some important blind spots?

Sachin: Most teams are watching the model very closely and almost everything around it hardly at all.

And I understand why. The model is the part the vendor demonstrates, so it is where the attention naturally goes.

But in my experience, the risk usually sits in the plumbing around it rather than in the model itself.

The blind spots I would name are the training data, which very few teams treat as the GxP input it really is; the life of the system after go-live, including drift; and the quiet, unofficial use of tools like ChatGPT and Copilot that is almost certainly already happening in the building.

But the biggest one, the one I keep coming back to, is decision integrity.

We still pour our energy into proving that the record is accurate.

The harder and more important question is whether the decision the AI helped shape was actually sound, and whether you could defend it.

That is where I expect the next wave of inspection findings to land.

“We’ve spent years proving the record is accurate. The harder question AI forces on us is whether the decision behind it was sound.” — Sachin Bhandari

Which AI risks are most organisations still missing?

Employees are already using tools like ChatGPT and Copilot to draft SOPs, summarise investigations, and support documentation, often without formal oversight. How serious is this “shadow AI” problem from a data-integrity perspective?

Sachin: It is a serious problem, and it is serious precisely because it is invisible.

Unvalidated AI is already sitting inside GxP documents in most organisations, and Quality usually cannot see it.Remember that data integrity is already among the most frequently cited categories in FDA drug GMP warning letters, mentioned in roughly 61 percent of those issued in 2021 by one European Pharmaceutical Review analysis [1], so shadow AI is quietly pouring untraceable content into the exact area inspectors scrutinise hardest.

An SOP or an investigation summary that a model drafted carries no attribution, no version, and no audit trail, yet it reads exactly as though a person sat down and wrote it, when they did not.The scenario that genuinely worries me is a regulator finding it before the company does, tucked inside a controlled document whose origin nobody can actually account for.

A ban is not the answer, and I say that having watched bans simply push the behaviour underground, where you cannot see it at all.The better path is to give people a sanctioned, governed tool for the low-risk drafting they are clearly already doing, and then to draw a very clear line around the things AI must never touch without review.

Picture an analyst dropping a messy investigation into a public chatbot, getting back a beautifully clean summary, and pasting it into the report.It reads wonderfully. But the source data has left the building, and not a single sentence can be traced back to where it came from. That is a data-integrity gap hiding comfortably inside good writing.

“Shadow AI is a data integrity gap hiding inside good writing. It reads beautifully, and nobody can say where a sentence came from.” — Sachin Bhandari

Does training data become a GxP record?

Training data rarely gets discussed through a data-integrity lens. When an AI model is trained on historical data, does that training data effectively become a GxP record? What happens if that data is later found to be flawed or biased?

Sachin: Yes, without hesitation.

If a model is supporting GxP decisions, the data it learned from is a GxP input, and it should be governed like one.

This is probably the most under-discussed risk in the field right now.

Governing it means actually knowing where it came from, knowing it is sound, knowing whether it genuinely represents what you are now using the model for, and being able to show all three if you are asked.If that data turns out later to be flawed or biased, you do not just have a problem going forward. You have one reaching back across every decision the model has touched since.

The logic is no different from a contaminated reference standard, where you have to assess the impact backwards rather than simply fixing things from today onward. The uncomfortable truth is that most firms could not, if you asked them this afternoon, produce the lineage of their training data, and that gap is a finding waiting to happen.

Picture a model that learned what a good batch looks like from years of history drawn mostly from one site. Put it to work across the whole network and it will quietly mark the other sites down. The model is not broken at all. The data was never representative in the first place, and nobody kept its lineage.

“A validated model trained on unvalidated data is a contaminated system, and the contamination is invisible until the inspection.” — Validating AI in GxP: A Practitioner’s Guide, Chapter 1

What risks come with AI embedded in commercial platforms?

Many organisations are beginning to rely on AI capabilities embedded within commercial platforms, even though they have limited visibility into how those models were trained, updated, or governed. From a data-integrity perspective, what risks does that create, and what evidence should organisations expect from vendors before trusting those outputs?

Sachin: The central risk is silent change.

The vendor updates the model somewhere behind the scenes, and your validated behaviour shifts underneath you without a single change record on your side.

So supplier oversight has to reach all the way to the model and not stop politely at the software, which in practice means your audit trail has to take account of theirs.

Before I would trust the output, I would want to see how the model was trained and validated, how and when it gets updated, whether you are told before a change lands, how it performs for your particular use rather than in general, and how the vendor will actually stand behind you during an inspection. Notification of change is something I would want written into the contract, not left to goodwill.

And the accountability does not travel with the model. The regulator holds you responsible for the decision, whoever’s model happened to produce it, so “the vendor takes care of that” was never going to be a defence.

The case I have seen play out is a QMS vendor quietly improving its deviation classifier in a routine release. Overnight, borderline events start being categorised differently. You validated the old behaviour, you are now running the new one, and you found out about it from a trend in your own data rather than from a change notice.

“The regulator holds you responsible for the decision, whoever’s model produced it. ‘The vendor takes care of that’ was never a defence.” — Sachin Bhandari

What will AI inspection readiness look like?

If a regulator challenged a decision that had been influenced by AI, what would an organisation need to demonstrate? If you had to build that evidence package yourself, what would absolutely need to be in it?

Sachin: They would need the whole story of the decision, end to end, told well enough that they could rebuild it without me standing in the room.

The pack would have to carry the intended use and the risk assessment, a clear statement of what was validated and the criteria you held it to, and then the decision record itself. By that I mean the input or prompt, the model and its version, the output, and the person who accepted it.

Alongside that, you would want the monitoring evidence showing the model was sitting inside its validated state at that particular moment, and the history of any changes or retraining since. And on top of all of it, the evidence of human oversight: who reviewed it, what they were actually checking, and that they had the standing to overrule it if they had needed to.

Put as simply as I can, I would want to show that a competent person, using a controlled system that was performing within limits we had defined, made an accountable decision, and that I can prove every single link in that chain. It is really not so different from a batch record, only for a decision instead of a product.

You have the inputs, the equipment in the shape of the model and its version, the in-spec evidence from your monitoring, and the signature at the end. Leave a gap in any one of those links and the decision becomes exactly as hard to defend as a batch record with a hole in it.

Do regulators need a fundamentally different inspection model?

Regulators are still largely applying existing frameworks to AI. Do you think that is sustainable, or are we heading towards a point where a fundamentally different inspection model is needed?

Sachin: For now, it is sustainable, and I actually think it is the right call, because the existing frameworks stretch a good deal further than people give them credit for.

ALCOA+, Annex 11, Part 11, GAMP 5, and ICH Q9, between them, cover most of what AI needs, provided you are willing to interpret them honestly rather than mechanically. The real gap is not in the principles. It is in the specific guidance, and that is already arriving through the draft Annex 22, the EU AI Act, and the work coming out of bodies like ISPE and NIST.

Where I think the current model will eventually start to strain is with systems that keep learning and changing while they are in production, because a once-a-year snapshot inspection suits a system that holds still, not one that is always quietly moving. My expectation is that we drift towards more continuous evidence and monitoring expectations, but inside the framework we already have rather than through some brand-new one bolted on beside it.

Think of it like inspecting a process that retunes itself every shift. A single annual photograph tells you very little, and the direction of travel is inspectors reading your monitoring evidence, not just leafing through your binders.

What is the industry still underestimating?

What aspect of AI and data integrity do you think the industry is still underestimating today?

Sachin: Decision integrity, and running close behind it, the fact that validation has become a continuous obligation rather than a one-off event.

We spent years getting data integrity genuinely right, and AI quietly lifts the bar to a harder question, which is whether the decision itself was sound, not merely whether the record of it was accurate. People still tend to treat AI validation as something you complete and then file away.

But the system does not hold still, so the obligation cannot either. And training data as a GxP record in its own right is still barely on anyone’s radar.

The thought that really brings it home for me is this: a perfectly accurate, fully attributable record of a poor AI recommendation will still sail through a data-integrity check, and it will still lead you to the wrong release. That gap, between a clean record and a sound decision, is exactly the space decision integrity exists to fill.

“A perfectly accurate, fully attributable record of a bad AI recommendation still passes a data integrity check, and still leads to a wrong release.” — Sachin Bhandari

What question should quality leaders be asking?

What question should quality leaders be asking about AI and data integrity that almost nobody is asking today?

Sachin: The question I would want them asking is this: if this AI were wrong, how would we even know, and how far back would the damage reach?

Almost everyone asks whether the model is accurate. Far fewer stop to ask about the spread of a quiet failure, or how long it might run before anybody actually noticed something was off. And the follow-up that hardly anyone asks at all is where AI is already sitting inside our GxP records without ever having been governed.

Most leaders assume the honest answer is nowhere, and in my experience the honest answer is almost never nowhere. If you sit a quality team down and ask them what the very first sign of drift would be, and who in the building would see it, the silence that usually follows is the real finding.

“Ask a quality team what the first sign of drift would be, and who would see it. The silence that follows is usually the real finding.” — Sachin Bhandari

What should quality leaders remember?

If quality leaders remember only one thing from this conversation about applying ALCOA+ to AI, what would you want it to be?

Sachin: If they hold on to nothing else, I would want it to be this: AI does not retire ALCOA+.

It raises the bar on what it takes to satisfy it, and it adds one new layer on top, which is that you now have to defend the decision and not just the record.

Said as plainly as I can manage, the principles are the same as they always were, the evidence is harder to assemble, and there is a new floor underneath the whole thing called decision integrity. The firms that come through this well will not be the ones with the most impressive models in the building.

They will be the ones who can still answer the oldest question in our field, “Can you prove it?”, for a system that thinks for itself. Twenty years ago, that question meant, “Can you prove this number is real?” Today it is growing into, “Can you prove this judgement was sound?”

The instinct behind it has not changed at all. The stakes have simply gone up.

“The firms that come through this well won’t have the fanciest models. They’ll be the ones who can still answer the oldest question in our field, ‘Can you prove it?’, for a system that thinks.” — Sachin Bhandari

From data integrity to decision integrity

The shift towards AI does not remove the need for accurate, attributable, contemporaneous, original, complete, consistent, enduring, and available records.

It expands the scope of what organisations must control. The model’s behaviour, its training data, its version history, the surrounding workflow, the commercial vendor, the human reviewer, and the final decision all become part of the evidence chain.

That means validation can no longer end at go-live. It must continue through monitoring, change control, retraining, supplier oversight, and periodic evaluation of whether the model remains fit for its intended use.

The central question for inspection readiness is no longer only whether an AI-generated record can be reconstructed. It is whether the organisation can prove that a controlled system, operating within defined limits, supported a sound decision that a competent person consciously accepted and owned.

If your organisation is working through what model drift, shadow AI, and decision integrity mean for its validated systems, BGO Software’s regulatory and compliance advisory team works at exactly this intersection. Get in touch to continue the conversation, and read the companion piece on translating ALCOA+ for AI for the data-integrity foundation underneath all of this.

About Sachin Bhandari

Sachin Bhandari is a Quality IT and digitalisation leader with more than 25 years in pharma, working at the intersection of CSV, CSA, and the validation of AI and ML systems in GxP environments.

He has led enterprise eQMS, paperless validation, and cloud quality programmes at global scale, and has taken AI-supported quality systems through health-authority inspection in live operations.

He is the founder of TrustBridge Compliance and the author of Validating AI in GxP: A Practitioner’s Guide. More at trustbridge-compliance.com.

Reference notes

Reference notes (for editorial verification, current as of June 2026). Note: EU GMP Annex 22, the revised Annex 11, and Chapter 4 are drafts published 7 July 2025 (consultation closed October 2025), with finalisation expected in 2026.

[1] Data integrity among the most frequently cited categories in FDA drug GMP warning letters — European Pharmaceutical Review, “FDA warning letters highlight data integrity issues” (61 percent of 2021 letters); see also RAPS on common warning-letter citations.

[2] Model drift, lifecycle risk, and AI governance — NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0), which addresses AI systems trained on data that changes over time. (Under revision per the July 2025 US AI Action Plan; cite as AI RMF 1.0.)

[3] AI supporting regulatory decisions, vendor oversight, and inspection readiness — FDA draft guidance “Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products” (January 2025), which also lists FDA’s “Artificial Intelligence in Drug Manufacturing” discussion paper.

[4] EU GMP draft Annex 22 (AI), Annex 11 and Chapter 4 — ECA Academy / GMP-Compliance: drafts released for comment and the European Commission consultation hub.

[5] EU AI Act, Regulation (EU) 2024/1689 (applying from 2 August 2026) — official EUR-Lex text and the European Commission regulatory framework overview.

[6] AI-specific validation guidance for GxP computerised systems — ISPE GAMP Guide: Artificial Intelligence (July 2025) and ISPE GAMP 5 Guide, Second Edition.

[7] Quality risk management and risk-based decision-making — ICH Q9(R1) Quality Risk Management (EMA), whose revision added a dedicated section on risk-based decision-making.

bgosoftware logo

BGO Software

BGO Software is a renowned IT company specializing in healthcare technology solutions. With over 15 years of industry experience, we offer comprehensive technology consulting, software development, and IT infrastructure services. Our focus on healthcare technology enables businesses to mitigate technological risks and accelerate growth, ultimately delivering enhanced value to patients.

link to the author’s linkedin profile

What’s your goal today?

Hire us to develop your
product or solution

Since 2008, BGO Software has been providing dedicated IT teams to Fortune
100 Pharmaceutical Corporations, Government and Healthcare Organisations, and educational institutions.

If you’re looking to flexibly increase capacity without hiring, check out:

On-Demand IT Talent Product Development as a Service

Get ahead of the curve
with tech leadership

We help startups, scale-ups & SMEs create cutting-edge healthcare products and solutions by providing them with the technical consultancy and support they need to break through.

If you’re looking to scope and validate your Health solution, check out:

Project CTO as a Service

See our Case Studies

Wonder what it takes to solve some of the toughest problems in Health (and how to come up with high-standard, innovative solutions)?

Have a look at our latest work in digital health:

Browse our case studies

Contact Us

We help healthcare companies worldwide get the value, speed, and scalability they need-without compromising on quality. You’ll be amazed of how within-reach top service finally is.

Have a project in mind?

Contact us
chat user icon

Hello!

Did you know that BGO Software is one of the only companies strictly specialising in digital health IT talent and tech leadership?

Our team has over 18 years of experience helping health startups, Fortune 100 enterprises, and governments deliver leading healthcare tech solutions.

If you want to explore your options, would you like to book a free consultation call today?

Yes

It’s a free, no-obligation, fact-finding opportunity. You’ll have a friendly chat with our team, ask any questions, and see how we could help in detail.