Ask ten PV teams if their AI tools are compliant, and you'll get ten confident yeses. Ask them to walk you through how they'd actually prove that to an inspector, and the confidence drops fast.
That gap is the part of the AI-in-PV conversation nobody wants to sit in. We talk a lot about what AI can do — flag signals faster, triage literature, cut down on manual case processing. We don't talk enough about what happens when someone asks you to show your work. "Audit trails" and "GxP-compliant infrastructure" show up in almost every vendor pitch, including ours, like they're a solved problem. They're not a checkbox. They're a set of specific, answerable questions, and inspectors are starting to ask them about AI-assisted decisions specifically, not just about the system as a whole.
AI adoption in PV is moving faster than the audit-readiness conversation is maturing. This is where I think that conversation needs to catch up.
Strip away the vague language and it comes down to three concrete questions.
Can you show your work? Not just that a signal got flagged or a literature hit got triaged, but why. What did the tool see, what did it weigh, what did a person confirm or override. This is where the difference between "AI-assisted" and "AI-decided" actually matters, and it's worth being precise about which one you're running. If a human is confirming every output before it moves forward, that's a very different audit story than a tool making a call on its own.
Who's accountable for what? Human-in-the-loop isn't a slogan, it's a boundary that needs to be drawn somewhere specific. Where does the tool's output stop and a qualified person's judgment start? Good PV teams already think this way about their people — RACI structures, defined sign-off points, clear ownership. AI tools need the same discipline. The tool doesn't get a seat at the table where accountability lives; a person does.
Can you reproduce it? If an inspector asks you to explain a decision from six months ago, can you actually do it? Traditional validated systems are static enough that this is usually straightforward. AI tools complicate it — model updates, retraining, version changes can all shift what the system would do with the same input today versus six months ago. If you can't reconstruct the conditions under which a decision was made, you can't defend it.
The guidance here is out there, but most of it reads like it was written for lawyers, not for the people running PV operations day to day. Here's the plain-language version.
The EU AI Act puts PV signal detection in the high-risk category. In practice, that means the burden of proof shifts. You're not just expected to have a good tool — you're expected to be able to demonstrate, on request, that it's being governed the way a high-risk system needs to be.
The joint EMA-FDA Good AI Practice principles and the EMA reflection papers land on a similar set of expectations even without the same legal weight behind them: documentation at the point of use, defined human oversight, and ongoing monitoring of how the system is performing, not just a one-time validation at launch.
FDA's draft guidance follows the same thread. None of these documents are asking for something exotic. They're asking for the same rigor PV teams already apply to everything else, extended to a new kind of tool.
There's a distinction that comes up constantly in regulated Salesforce implementations: did the feature work as built (UAT), versus can you trust the system's ongoing output in a GxP context (validation). With static software, those two questions mostly converge once you're past go-live. With AI, they don't.
Take a literature surveillance tool that passed validation at launch. If it gets retuned or retrained six months later, does its triage logic still behave the way it did on day one? Maybe. Maybe not. The system isn't static anymore, and "we validated it once" stops being a sufficient answer. Ongoing validation for AI-enabled workflows has to account for a system that can genuinely change underneath you, which is a different governance problem than most PV teams have had to solve for before.
A few things worth having in place before an inspector asks, not after:
None of this is about slowing AI adoption down. It's about being able to stand behind it.
AI doesn't remove the need for PV judgment. If anything, audit-readiness is the clearest proof of that. A system you can't explain isn't one you can defend, no matter how clean its output looks on a dashboard. The teams that get real value out of AI in PV over the long run won't be the ones who adopted fastest — they'll be the ones who could always answer the question, "how do you know it's right?"