CMC has a knowledge-continuity problem, not a data problem
Connected data and traceable decisions are becoming vital in a new era of pharmaceutical development, AI, and advanced therapies
17 Sept 2026

Ken Forman, Lead Product Manager at IDBS (left) and Mark Buswell, Director of AMB Engineering and Technology and a former Vice President at GSK (right)
Chemistry, manufacturing and controls (CMC) generates enormous quantities of scientific data as a drug candidate moves towards a reproducible, scalable and ultimately manufacturable medicine. Increasingly, the challenge is no longer simply whether CMC data has been digitalized, but whether the scientific knowledge, context and decision-making rationale behind that data remain connected as a product moves through development.
From establishing the quality target product profile (QTPP), through process development and characterization, scale-up, tech transfer, manufacturing and eventually Module 3, each stage adds new experimental evidence, assumptions, and decisions that teams downstream may need to reexamine/retrace/revisit. Electronic laboratory notebooks (ELNs), laboratory information management system (LIMS) platforms, and other digital systems can improve access to individual records, but they do not automatically preserve the relationship between a result, the evidence behind it, and the scientific decision it informed.
“Truly digitally mature companies have the linkages between the data and the decisions they have informed,” says Ken Forman, Lead Product Manager at IDBS. “A company's CMC system becomes not just a database, but a knowledge base.”
Tracing the decision, not just the data
Mark Buswell, Director of AMB Engineering and Technology and a former Vice President at GSK, shares the example of a one-hour process hold established experimentally during development.
The original experiment may have demonstrated that the process can safely remain at that stage for an hour, with this value subsequently carried into development reports, process documentation, and, years later, a regulatory filing. The source data may have been recorded compliantly, but as results move into reports and other documents, the link to the original record can be lost. “Ten years down the line, a regulatory team can be left working backwards through tables and reports to demonstrate where each figure originated,” he says.
Even when the original experiment is found, it only answers part of the question. “There’s traceability of data, but there’s also lineage of decisions,” explains Forman. “It’s not just one hour. It’s ‘someone decided that one hour is the number we will use for this control’.”
“What happens is the data gets handed forward, but what arrives is often just a collection of results,” he continues. “The next team is then left to reconstruct the logic that turned those results into a decision.”
That becomes harder still where the knowledge was tacit. Development decisions can be shaped by conversations between scientists, engineers and manufacturing teams who may have changed roles or left the organization long before the product reaches launch.
For mature CMC organizations, preserving knowledge continuity therefore means connecting, storing, and enabling information retrieval beyond source data and final results. Experimental parameters, product and process specifications, the decisions drawn from those results and the reasoning behind them all need to remain linked, alongside any subsequent changes those decisions trigger. It is this decision lineage that enables later teams to understand not only what was decided, but why.
There's traceability of data, but there's also lineage of decisions.
Ken Forman, Lead Product Manager, IDBS
Where knowledge continuity starts to break
These gaps can become apparent before formal tech transfer begins. During scale-up, teams can’t assume that process parameters and operating ranges established at development scale will translate directly to larger equipment and manufacturing conditions. They need the supporting experimental evidence, modeling and simulation results and assumptions. This body of data is then used to understand which parameters are scale-dependent, assess how changes in equipment, process volumes and manufacturing conditions could affect product quality and determine what must be adapted, verified or controlled at the next scale. When that context is difficult to recover, teams may need to revisit earlier studies or repeat work before they can scale the process with confidence.
The same problem becomes particularly visible during tech transfer. Development and manufacturing have inherently different requirements, and the systems supporting them frequently reflect that. “If you think about R&D systems, they’re designed for flexibility,” says Buswell. “Whereas in manufacturing, they’re designed to remove flexibility and really lock things down.”
As a result, development and manufacturing systems may be configured so differently that even organizations using the same LIMS platform struggle to transfer methods or data directly between them. “In some cases, teams simply key it in again because integrating the systems is more difficult than manually re-entering the information,” says Buswell.
The challenge becomes greater again when external manufacturing partners are involved. Direct system-to-system connectivity between sponsors and contract development and manufacturing organizations (CDMOs) is often unrealistic because of security, IT, and data-separation requirements. What matters, therefore, is whether the information that does move between organizations carries enough context with it.
Forman describes a receiving manufacturing team questioning why, for example, a pH set point has been defined within a particular range.
“Why? What is the defence of those numbers? What experiments did you run? What studies were performed? How did you come to that conclusion?”
If answering those questions requires the sponsor to search disconnected systems or reconstruct historical reasoning, the consequences extend beyond inefficient data management. Decisions can be delayed, receiving organizations or regulatory agencies left waiting and, in some cases, go-live or process performance qualification put at risk.
“There’s a joke in the industry that if you're in manufacturing and a problem arises, it can often be quicker just to repeat the original experimentation than track down the work that answered the same question 15 years earlier,” Buswell adds.
Increasingly complex therapies are making these gaps in knowledge continuity more apparent. Advanced modalities can involve additional scientific functions, manufacturing stages, logistics and external partners, increasing the number of points at which knowledge must move without losing its context. In autologous cell therapies, for example, logistics, scheduling, chain of custody, and even the patient can become part of the manufacturing process itself, further extending the information that needs to remain connected.
Building submission readiness through development
For mature CMC organizations, the goal is not simply to apply more digital tools or achieve perfectly clean data, but to preserve the connection between source data, how it is interpreted and how it is subsequently used. Buswell likens the principle to keeping a data element “hyperlinked” back through the development history to the validated experiment from which it originated.
This enables submission readiness to start much earlier than submission preparation. Capturing and reusing scientific information, including risk assessments, critical quality attributes (CQAs), and critical process parameters (CPPs), helps preserve the rationale behind key decisions. This requires more than keeping the final value. Organizations need to maintain relationships between product requirements, risk assessments, experimental evidence, CQAs, CPPs and the decisions they support. Unfortunately many organizations are still using cumbersome spreadsheets to try and maintain these relationships. Submission readiness then becomes an outcome of the development process, rather than a retrospective activity requiring investigation and reconstruction.
Putting connected knowledge to work
Preserving continuity also creates opportunities beyond tech transfer and regulatory preparation, particularly as CMC makes greater use of in silico modeling, machine learning, and AI.
Done well, this can create a much tighter loop between data, experimentation, and decision-making. Automated systems can trigger follow-up experiments, connect wet-lab work with in silico modeling, and increasingly support monitoring, root-cause analysis, AI and the early identification of emerging process issues.
Buswell argues that CMC organizations have sometimes delayed advanced analytics by pursuing the “holy grail of 100% structured, pure clean data.” Data quality remains important, but he sees it as a continuous process rather than a hurdle that must be completely cleared before organizations can extract additional value.
“The focus isn't so much about data quality as opposed to data usability; what is usable data for decision making when you bring AI into the equation?” asks Forman.
For AI and modeling, this shifts the emphasis from simply having large quantities of digital information to having information that can be interpreted in the context in which it was generated. The same links that help a manufacturing team understand why a process parameter was chosen, or a regulatory team defend the evidence behind it, also provide the context needed to use CMC data more effectively in advanced analytical approaches.
You should see more innovative medicines, more innovative formulations, more efficient routes delivered faster.
Mark Buswell, Director of AMB Engineering and Technology and a former Vice President at GSK
The objective is therefore not simply to create a more digital CMC organization. It is to create one in which scientific knowledge persists, remaining usable as the people, systems, and environments around it change.
For CMC leaders, the starting point is to identify where high-value knowledge is most likely to become detached from its evidence and rationale. This means prioritizing the lifecycle transitions and decisions that matter most for development, scale-up, tech transfer, manufacturing and submission, rather than attempting to connect every data source at once. “It’s no longer a data quality or data volume problem,” Forman concludes. “There’s lots of data, but it’s the continuity of the data and the knowledge around the data – having that available for the decisions you need to make and being able to defend them in a regulatory capacity.”