SPEAKERS:
Kris Bruynseels, VP & Global Head Integrated Insights, UCB
Gary McAllister, CTO Healthcare & Public Sector, Dell Technologies
Saibal Mukherjee, Director Global Data Digital & Technology Legal, Takeda (Moderator)
KEY TAKEAWAYS:
- AI compressed UCB's patient journey synthesis from eight to ten weeks down to hours
- "Laundering risk" describes AI outputs that reach leadership as validated strategy without genuine review
- Relabeling "AI oversight" as "expert review" measurably shifted teams from compliance posture to active judgment
- Workflow correction rates outperform model accuracy scores as the governing health metric
- Small-dataset AI models validated on narrow populations are entering pharma procurement without adequate regulatory scrutiny
The number that stopped the room at Pharma 2026 was not a trial readout or a regulatory milestone. It was eight to ten weeks compressed to hours. That is how long patient journey synthesis used to take at UCB. It now takes, on a slow day, until the following morning. "It used to take us, what was it, eight to ten weeks to get a synthesis of a patient journey. We now do that in hours, perhaps max a day. So the speed is there," said Kris Bruynseels, VP & Global Head Integrated Insights, UCB. "But if you then think about patients or even HCPs, I think they're less interested or concerned about the speed at which we generate insights, but more on: where does this come from? How reliable is it? What is the context behind it? How much can we trust it?"
That gap between what AI delivers and what downstream stakeholders actually need, is the structural problem the session spent an hour dissecting. Saibal Mukherjee, Director Global Data Digital & Technology Legal at Takeda and session moderator, framed the challenge through an aviation parallel: cruise control has governed 90% of commercial flight time since the 1970s, yet no airline has automated takeoff and landing. Pharma has not yet drawn the equivalent lines in its AI workflows, and the consequences of that omission are becoming visible.
Speed Creates Liability Before It Creates Value
The industry has largely treated AI's speed advantage as an unambiguous win. It is not. The same throughput that eliminates eight weeks of analyst work creates a structural condition where the volume of AI outputs exceeds any organization's capacity to meaningfully review them. When volume outpaces review, a specific failure mode emerges.
Bruynseels named it directly: "There's something in my field that we call a laundering risk, which is AI produces an output, insights teams would package it, send it to the brand teams or whoever is a consumer of it as a strategic recommendation. But there's no one that has really done an in-depth review of what came out. How reliable is it? What does it mean? What will we do with it?"
Laundering risk is not a technology deficiency. It is an organizational incentive problem. Teams optimized for speed have no structural reason to slow down for validation, doing so reads as friction against the very efficiency gains AI was procured to deliver. The acceleration and the exposure are the same phenomenon.
Gary McAllister, CTO Healthcare & Public Sector at Dell Technologies, located the consequence in legal architecture: "There is an IBM quote that I always use: you can't make a machine accountable for any human interaction. So the challenge we all have is that in a world where we're trying to create autonomous decision support tools, diagnostic tools, or even autonomous prescribing, who becomes accountable for something when it goes wrong? And that's the real question. And that's where humans in the loop have to be well established." When a laundered output drives a decision that turns out to be wrong, accountability defaults to people who never examined the underlying evidence, a position no legal framework currently resolves.
Mukherjee staked the commercial ceiling: "Commercial decision making has, I think, the highest risk, that decision making could lead to discussions towards off-label or any other area which could be very high risk." An AI-informed commercial recommendation that no human reviewed before reaching a brand team is not just an accuracy risk. It is a regulatory exposure.
The failure mode is real, but so is the counter-evidence. Bruynseels described a rare indication where an AI model detected an unusual delay between symptom onset and specialist referral, a signal that had always existed in the data but that prior analytical methods lacked the scale to surface. The signal alone was insufficient. "It wasn't until we had our MSLs connect with the clinical community to understand: is this a data anomaly, is it real, that made the decision that AI suggested actionable and meaningful." Same AI output, same underlying data. The difference between a potential artifact and a deployable clinical insight was a single, deliberately designed human validation step. That is what laundering risk forfeits.
Oversight Designed as Compliance Produces Compliance Behavior
The standard organizational response to AI risk is governance: councils, review layers, defined thresholds. Most of these structures produce the behavior they are designed to prevent. When people are assigned to "AI oversight," they process outputs. They do not evaluate them.
Mukherjee argued that governance frameworks, when constructed well, actually dissolve a more intractable problem: "Human oversight has enabled companies to break the silos, because AI governance or AI roadmaps cannot be done by a single function. It has to be done in consonance with all functions working together. Whether you call it Architecture Council or AI Council, that's the way to go in terms of crafting the governance to the workflows, to the product development, to the use cases." Cross-functional ownership is not incidental to good AI governance. It is the mechanism by which governance becomes operational rather than ornamental.
The more immediate intervention Bruynseels offered was terminological, and its effects were not cosmetic. "Initially we had what we called AI oversight; flipped it to expert review. And it is a small difference, but it moves it from a person controlling the heavy lifting that AI does and checking it for errors, to an expert review based on human capacity, context, knowledge. So the human judgment becomes more of the primary output than it is the AI model." The label change shifted the cognitive frame of the people doing the work. Oversight is a supervisory role. Expert review is a primary contribution. People perform the role they believe they have been given.
The structural complement to that reframe is a triage logic that concentrates expert judgment where it matters rather than spreading it uniformly across all outputs. Bruynseels described a working prototype: "Something that we've been working with is a relatively easy consequence ladder, two dimensions: what is our confidence in the models that we use, and what is the impact of making a wrong decision? If confidence is low and impact is high, I think that's where there's no debate. You need to have that human that puts in a total review, almost like an expert review, before you take it to any next step of decision making." Currently in testing at UCB, the framework addresses the volume problem directly. If every output receives the same level of scrutiny, no output receives genuine scrutiny. Risk-tiered intervention allows organizations to direct expert judgment to the decisions where a wrong answer causes real damage.
McAllister noted that mature AI platforms are already beginning to embed this logic at the tool level, providing anomaly detection and confidence scoring alongside outputs so clinicians and analysts can identify when override is warranted. The structural design of the technology and the organizational design of the review process are two halves of the same trust architecture.
Discover more on this topic at Pharma Commercial Data & Tech Europe 2026 (4-5 November, London) Europe’s collaborative home for data and tech pioneers. Visit the website here.
Small-Dataset Models Are the Unpriced Risk in Procurement
The oversight conversation inside most pharma organizations focuses on internally developed models. The more immediate exposure may be at the procurement desk.
McAllister identified a design philosophy problem that cuts across the vendor market: "I think one of the problems we have at the moment is AI tools are developed with the happy path. So the AI tool always expects a perfect outcome on the decision that it's made and doesn't necessarily understand the variation and the complexities of human condition. So we do need to get to a point where we've got AI that has variation and anomaly detection built within it, so that the AI can flag human intervention within the process." Tools optimized for ideal conditions are not just suboptimal in complex clinical settings. They are structurally blind to the cases where intervention is most needed.
What that blindness produces in practice is documented. McAllister cited UK academic research: "There are academic papers in the UK around a model that was trained for people with bowel disease on an elderly care ward, and the outputs that the model generated — because it was hallucinating — were getting people's sex wrong and their date of births wrong and giving people bowel conditions that they didn't even have." This was not a cybersecurity breach or a data pipeline failure. It was a model working exactly as designed, on a population it was never trained to handle.
The market entry path for these tools is the underlying problem. "There's a lot of small dataset AI models that have been created which in very strict circumstances provide you with high probabilistic outcomes, but when you then apply them to a broader demographic, they become highly unsafe — and they're the ones that we need to weed out. They're not going through appropriate levels of regulation or registration," McAllister cautioned. In the UK, Class 1 medical device self-certification creates a pathway that validates narrow-condition performance without assessing demographic generalizability. The model passes. The patient population it encounters in production does not match the population it was trained on. For pharma procurement teams, this is a due-diligence gap that current governance frameworks do not close.
The Metric That Reveals Whether Your Governance Is Real
McAllister identified the behavioral endpoint toward which inadequate oversight is trending: "The main worry is you get AI fatigue, and you end up in a place where AI becomes the only thing that makes a decision in your organization. That's a problem, and it isn't just a healthcare problem. It's a problem across every industry. We're getting slightly AI fatigued, and that means we're making poor decisions because we're not doing our own research." AI fatigue is not complacency. It is the rational response to review workloads that have outpaced human capacity, the same dynamic laundering risk creates, only further downstream.
What the panel's evidence collectively reveals is an inverse relationship: organizations deploying AI fastest are the ones most structurally exposed to laundering risk, because their output volume overwhelms review capacity first. AI governance maturity does not automatically track AI deployment velocity. It can move in the opposite direction unless organizations explicitly decouple review rigor from output volume. The consequence ladder and the expert review reframe are both mechanisms for that decoupling, but most organizations have not yet recognized they need it, because they are still measuring AI performance at the model level rather than at the workflow level.
Bruynseels proposed the measurement shift that closes this gap: "In an insights space, I'm a very strong believer of measuring the entire workflow — human-AI as a whole — how well is that performing? And within that, I try to look at correction rates. How often have we actually intervened on an AI output, and why? Was it the data? Was it context? That correction rate gives me a far better signal on how good we are and how confident we should be — and if we do it well, we feed back the true training data to make the system smarter." Correction rates communicate governance health in terms a brand leader with fifteen years of market experience can interpret. They also generate the feedback loop that improves the underlying model, converting human review from a cost center into a capability investment.
Mukherjee placed this in the only governance structure that makes it durable: "The change is meaningful if it comes from the top, from the top leadership, because then it's easier to adopt across the whole organization. It's not something which is owned by regulatory or compliance or legal. I think it's an organizational change we are looking at." Correction rate reporting belongs on the same dashboard as pipeline velocity and commercial readiness, not in a governance annex that leadership reviews quarterly.
The question every leadership team should be able to answer is simple: how often did human reviewers intervene in AI outputs last quarter, and what did those interventions change? If the answer requires a conversation with the data science team, the governance infrastructure is producing oversight theater. If the answer is immediately available and tracked against targets, the organization has begun to convert AI speed into something more durable than throughput. Speed that cannot be interrogated is not an asset. In pharma, it is a liability waiting for a decision.
To get you highlights of Pharma 2026 faster, we are using generative AI technology to summarise the transcripts of the sessions. If you have any feedback about the summary, please contact lucy.fisher@thomsonreuters.com.
Discover more on this topic at Pharma Commercial Data & Tech Europe 2026 (4-5 November, London) Europe’s collaborative home for data and tech pioneers. Visit the website here.