Esperimental Design Cpmes Before Technology
Scientific Perspective

A practical framework for human-relevant research
New Approach Methodologies (NAMs) are giving biomedical researchers an unprecedented range of experimental systems: human organoids, microphysiological systems and organs-on-chip, advanced cell models, computational approaches, artificial intelligence, and increasingly sophisticated combinations of these technologies.
But more tools do not automatically produce better evidence.
If anything, the greater the technological choice, the more experimental design matters.
During my years in academic research, I repeatedly saw the same pattern. A laboratory had an established mouse model, a protocol that worked, experienced personnel, and years of historical data. A new target or hypothesis emerged, and the almost automatic next question became:
“What happens if we test it in the mouse?” The problem was not the mouse. The problem was allowing the model to become the starting point rather than the scientific question.
There is also an important distinction between exploration and what scientists sometimes informally call “fishing.” Exploratory research is essential; it can reveal unexpected biology and generate entirely new hypotheses. The difficulty begins when many endpoints are collected without defining in advance what result would support or challenge the hypothesis, and an interesting signal discovered afterward is then treated as though it had been the intended test all along.
Good experimental design makes that distinction explicit. Frameworks such as ARRIVE emphasise prospective attention to outcomes, experimental units, sample size, bias control, and analysis so that the resulting evidence can be interpreted rigorously [22].
As NAMs expand the experimental toolbox, this discipline becomes more important, not less.
Otherwise, we risk simply replacing mouse-first science with organoid-first, chip-first, or AI-first science.
The technology changes.
The experimental-design challenge does not.
Start with the decision, not the model
Before asking: “Which model should I use?” I believe we should ask: “What decision does this experiment need to inform?”
That apparently small shift in language forces us to identify the uncertainty we are trying to reduce.
A practical sequence then follows:
Scientific question –> hypothesis –> biological context –> fit-for-purpose model –> endpoints and controls –> validation –> decision –> next experiment.
This is the framework illustrated in Figure 1 and operationalized in Table 1.

The biological-context step is particularly important.
Which cells need to be present?
Does human genetic background matter?
Is tissue architecture relevant?
Are immune interactions, stromal components, metabolism, mechanical forces, flow, chronic exposure, or tissue–tissue interfaces central to the mechanism?
Does the question require whole-organism pharmacokinetics, or can the uncertainty be resolved at the cellular or tissue level?
Only after these questions are considered should the model be selected.
And the best model is not necessarily the most sophisticated one.
A simple 2D assay may be entirely appropriate for establishing target engagement or a concentration–response relationship. An organoid may become necessary when tissue organization, maturation, or patient heterogeneity matters. A microphysiological system may add value when flow, repeated exposure, mechanical forces, or multicellular interfaces are essential. A computational model may be highly informative within a sufficiently characterized domain. And an animal study may remain the correct experiment when the unresolved question genuinely requires integrated physiology.
The principle is therefore not:
Choose the most human-like model available.
It is:
Choose the least complex fit-for-purpose system that captures the biology required to answer the question with sufficient confidence.
That is a fit-for-purpose decision.
The FDA’s March 2026 draft guidance on NAMs takes a similar view, placing validation within an intended use rather than treating a methodology as universally validated [1]. The broader implication matters: model credibility depends on what the model is expected to do, under what conditions, and for which decision.
A practical academic example
Consider an academic laboratory that identifies a promising anti-fibrotic target through transcriptomic analysis.
A conventional pathway might move relatively quickly into an established mouse fibrosis model, administer an inhibitor, and evaluate histology, gene expression, serum markers, and other outcomes.
Is the target expressed in the relevant human cell population?
Is the observed effect direct or secondary?
Is the active concentration biologically realistic?
Does the mechanism depend on interactions that differ between humans and the animal model?
And which endpoint actually determines whether the hypothesis should advance?
A question-first, NAM-informed approach would not necessarily remove the animal experiment. It would change when and why it is used.
The laboratory might first interrogate human datasets and existing evidence to establish target expression, pathway direction, disease stage, and the relevant biological population.
A relatively simple human-cell system could then test perturbation, concentration–response relationships, and mechanism-linked endpoints.
If tissue organization or heterogeneous cellular interactions prove important, the scientific question may justify progression to a 3D model or organoid.
If flow, tissue interfaces, repeated exposure, immune–stromal interactions, or mechanical forces are critical, an organ-on-chip or other microphysiological system may become appropriate.
The point is not to create a mandatory technological ladder.
It is to increase complexity because the scientific question requires it.
Research using Liver-Chip systems illustrates this principle. Jang and colleagues used multicellular liver systems under physiological flow to examine human and cross-species toxicities and showed that relevant drug responses can diverge between species [11].
The significance of that work is not simply that a chip is technologically sophisticated.
It is that the additional biology incorporated into the model addresses a specific translational uncertainty.
At every step, the primary endpoint and the decision criterion should already be visible:
What result allows us to advance?
What result requires refinement?
What result stops the hypothesis?
What uncertainty remains afterward?
The experiment should not simply generate information.
It should reduce uncertainty.
Validation is part of experimental design
The growing attention to NAM validation is another reason model selection cannot be separated from experimental design.
Validation should not be treated as a generic certificate attached to a technology.
At the research level, it can begin with far more practical questions:
Does the system work technically?
Does it reproduce the biology required for this question?
Is its performance sufficient for the decision we intend to make?
Answering these questions means considering cell identity, functional phenotype, positive and negative controls, donor or cell-line variability, independent experimental runs, technical stability, benchmark compounds, exposure, and, where appropriate, transferability between operators or laboratories.
The Liver-Chip performance study by Ewart and colleagues provides a useful example of decision-oriented benchmarking. In a blinded evaluation of 27 known hepatotoxic and non-toxic drugs, the system achieved 87% sensitivity and 100% specificity under the study conditions [12].
The lesson is not that every liver-safety question now requires a Liver-Chip.
It is that performance was evaluated prospectively against known outcomes in a defined context of use.
Organoids raise similar considerations. They can capture human genetics, tissue-specific biology, and patient heterogeneity, but reproducibility may depend on protocol, stem-cell state, maturation, batch, and site. Cross-site cortical organoid studies show both encouraging reproducibility and measurable sources of variation, reinforcing the need for characterization rather than an assumption of biological fidelity [18].
Computational approaches demand the same discipline. The FDA’s acceptance of the first in silico drug-development tool into the ISTAND qualification pathway for predicting drug-induced liver injury is explicitly tied to a specified context of use and a broader weight-of-evidence strategy [14].
Different technologies therefore require different validation strategies, as summarized in Table 2.

No platform is universally superior.
Human relevance is not synonymous with predictive validity.
And complexity is not automatically fidelity.
Where NAM education can make a difference
This is also where I see an important educational opportunity for NAMina.
Academic laboratories do not necessarily need another catalogue of sophisticated technologies. They need a practical way to convert a scientific question into an experimental evidence pathway.
A researcher may already know that organoids, organs-on-chip, computational models, and advanced co-culture systems exist.
The harder questions are:
When should each be used?
What biology must the model reproduce?
What controls are required?
How should the system be benchmarked?
What constitutes sufficient validation for the intended conclusion?
When should complexity increase?
And what evidence should trigger the next experiment?
This is where education in NAMs becomes inseparable from education in experimental design.
The repeatable discipline is simple:
Define the decision. Identify the biology that matters.
Choose a fit-for-purpose model. Prespecify the endpoint.
Build validation into the experiment. Let the result determine the next step.
NIH’s Complement-ARIE initiative reflects a similar view of the field, linking technology development with standardization, validation, data infrastructure, deployment, and workforce training [4]. Its Validation and Qualification Network is intended to advance standardized reporting, quality systems, and validation or qualification frameworks [25].
In other words, successful NAM adoption is not simply a technology-transfer challenge.
It is a scientific-capability challenge.
And that distinction matters.
Teaching researchers that a new technology exists is useful.
Teaching them when to use it, why to use it, how to validate it, and what decision it should support is far more consequential.
Measure decisions, not experiments
If question-first experimental design is valuable, its success should also be measurable.
Simply counting how many NAM experiments were performed is not enough.
More meaningful indicators include the proportion of studies with a defined hypothesis and primary endpoint before experimentation; reproducibility across independent runs, batches, donors, or sites; performance against reference compounds or known perturbations; and the proportion of experiments that actually change an advance, refine, or stop decision.
I would also argue for measuring:
time-to-decision and cost-per-decision rather than only cost-per-experiment.
A more expensive model can be highly efficient if it resolves a critical uncertainty early.
A cheap experiment can ultimately be costly if it produces large amounts of data without changing what happens next.
The same principle applies to the impact of NAMs on animal use.
Success should not be judged only by counting one-for-one replacements.
A human-relevant experiment may instead refine an animal study, prevent an unnecessary one, identify the appropriate dose or endpoint, select the most informative candidates for in-vivo evaluation, or make a later animal experiment substantially more informative.
NAMs can therefore change when an animal experiment is necessary and what question it is being asked to resolve, rather than simply replacing it.
These potential measures are summarized in Table 3.

From model selection to evidence design
The future of NAMs should not be reduced to a competition between the mouse, the organoid, the chip, and the algorithm.
Each can be valuable.
Each can also be used poorly.
The more useful question is:
What combination of evidence gives us the greatest confidence to make the next scientific decision?
For some questions, the answer may begin with computation.
For others, a simple human-cell assay.
In another program, tissue architecture, multicellular interaction, or flow may be essential from the outset.
And for genuinely organism-level questions, an animal model may remain indispensable.
There should be no ideological hierarchy of models.
There should be a hierarchy of questions, uncertainties, and evidence.
That distinction matters as NAMs move from specialist technologies toward broader adoption in academic research, drug development, and regulatory science.
The goal is not to replace one default model with another.
It is to replace reflexive model choice with deliberate evidence design.
And that begins before the first experiment is performed.
Dr Paola Dama, PhD
Founder & CEO, NAMina Bio | Research Fellow, Sussex University
Figure 1. Experimental Design Comes Before Technology

Figure legend. A question-first evidence pathway for translational research and NAM selection. Experimental design begins with the scientific question, hypothesis, and biological context before a fit-for-purpose model is selected. Endpoints, controls, and validation criteria are established before interpretation. The resulting evidence supports a decision and determines the next experimental iteration. The framework does not rank technologies; model complexity should increase only when required to resolve the relevant scientific uncertainty.
Bracketed numbers correspond to the reference list in the accompanying article.
References
Reference numbering is retained to correspond with the accompanying figure and tables.
[1] U.S. Food and Drug Administration. General Considerations for the Use of New Approach Methodologies in Drug Development. Draft Level 1 Guidance. March 2026.
[4] National Institutes of Health Common Fund. Complement Animal Research In Experimentation (Complement-ARIE) Program.
[11] Jang K-J, Otieno MA, Ronxhi J, et al. Reproducing human and cross-species drug toxicities using a Liver-Chip. Science Translational Medicine. 2019;11(517):eaax5516. doi:10.1126/scitranslmed.aax5516.
[12] Ewart L, Apostolou A, Briggs SA, et al. Performance assessment and economic analysis of a human Liver-Chip for predictive toxicology. Communications Medicine. 2022;2:154.
[14] U.S. Food and Drug Administration. FDA Accepts First In Silico Drug Development Tool Under ISTAND Program to Help Predict Drug-Induced Liver Injury.
[18] Glass MR, Waxman EA, Yamashita S, et al. Cross-site reproducibility of human cortical organoids reveals consistent cell type composition and architecture. Stem Cell Reports. 2024;19(9):1351–1367. doi:10.1016/j.stemcr.2024.07.008.
[22] Percie du Sert N, Hurst V, Ahluwalia A, et al. The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research. PLOS Biology. 2020;18(7):e3000410. doi:10.1371/journal.pbio.3000410.
[25] National Institutes of Health Common Fund. Complement-ARIE Validation and Qualification Network (VQN).



![[Sliding Doors] The Golden Cage or the Courage to Build Freely?](https://static.wixstatic.com/media/b335f1_9352bb70137240c4ab71a3129c584931~mv2.png/v1/fill/w_980,h_1225,al_c,q_90,usm_0.66_1.00_0.01,enc_avif,quality_auto/b335f1_9352bb70137240c4ab71a3129c584931~mv2.png)


Comments