How to Choose the Right Variables for Reaction Optimization
A practical framework for choosing factors, ranges and responses that produce a useful reaction optimisation campaign.
By AutoRxn editorial team
The short version: include a variable because it could change the decision you make—not simply because it can be changed. Start with the response you need, identify the controllable factors most likely to affect it, and keep ranges chemically meaningful, safe and executable.
Choosing variables is one of the highest-leverage decisions in reaction optimisation. It defines what the campaign can learn, which interactions it can expose and how much experimental budget is needed. A wide search space can feel thorough while making the results harder to interpret. A narrow one can converge neatly on the wrong region.
The aim is not to list every possible condition. It is to define the smallest search space that can answer the development question.
Begin with the decision
Before choosing factors, write down what the campaign must decide. For example:
- identify viable catalyst and ligand families for a new transformation;
- increase assay yield without increasing a critical impurity;
- find conditions that remain reliable across several substrates;
- define a robust operating region for scale-up;
- reduce catalyst loading while holding conversion and selectivity within specification.
Each objective implies a different experiment. A catalyst-screening campaign needs meaningful categorical choices. A robustness study needs ranges around a candidate operating point and deliberate attention to noise factors. A scale-up study may need mixing, dosing or heat-transfer variables that were irrelevant during discovery.
Define the response at the same time. “Best reaction” is not measurable. Assay yield, isolated yield, conversion, regioselectivity, enantiomeric excess, impurity level, cycle time and process mass intensity answer different questions. If several responses matter, record them separately even if the optimiser ultimately uses a combined objective.
Separate factors from context
A factor is something the campaign deliberately varies. A response is an outcome it measures. A context variable records how the experiment was run but is not intentionally varied.
This distinction prevents accidental conclusions. Operator, reagent lot, plate position, analytical method and execution date may explain variation, but they are not automatically optimisation factors. Record them as metadata so that batch effects and drift can be investigated.
| Role | Reaction example | How to handle it |
|---|---|---|
| Controllable factor | Temperature set point | Define a feasible range and setting precision |
| Categorical factor | Ligand identity | List available, chemically credible choices |
| Response | Assay yield | Specify method, units and direction of improvement |
| Constraint | Pressure limit | Prevent suggestions outside the safe region |
| Context variable | Reagent lot | Record consistently; inspect for hidden effects |
Use chemistry to build the candidate list
Mechanistic understanding, precedent and preliminary observations should determine which variables enter the first campaign. Ask where the reaction network is likely to be sensitive:
- Which step controls rate or selectivity?
- Which species may change aggregation, speciation or catalyst resting state?
- Could solubility or mass transfer be limiting?
- Which components influence competing pathways or catalyst decomposition?
- Are temperature and time interchangeable, or might they change the product distribution differently?
- Which conditions are likely to become impractical at larger scale?
For a metal-catalysed coupling, catalyst precursor, ligand, base and solvent may be strongly coupled. Treating one as important and fixing the others too early can hide a viable combination. Conversely, adding stirrer speed, vessel type and atmosphere to every small-scale campaign may spend runs on factors that are already controlled adequately.
Classify variables correctly
The variable type affects both the experimental design and the model.
Continuous variables
Temperature, time, concentration, flow rate and catalyst loading are usually continuous within physical limits. Record the achievable precision: a nominal range of 40–41 °C is not meaningful if the equipment controls temperature only to ±2 °C.
Integer variables
Some settings are discrete but ordered, such as the number of equivalents expressed in fixed increments, residence-time steps imposed by hardware, or the number of reactor cycles. Do not treat them as continuous if intermediate values cannot be executed.
Categorical variables
Solvent, base, catalyst, ligand and reactor type are labels rather than points on a numeric scale. Categories should represent actual experimental choices. Arbitrary numbering—such as solvent 1, 2 and 3—must not imply that solvent 2 lies halfway between solvents 1 and 3.
Descriptors can be useful when they are available and relevant, but a poorly justified descriptor set can add apparent precision without adding chemical information.
Conditional variables
Some factors exist only under certain choices: ligand loading is irrelevant for a ligand-free catalyst, and light intensity is irrelevant to a thermal branch. Make these dependencies explicit rather than filling the table with misleading zeroes.
Choose ranges that can teach you something
A useful range is broad enough to produce a measurable response but narrow enough to remain safe, feasible and scientifically connected.
For each numeric factor, document:
- the hard physical or safety limits;
- the practical operating limits of the equipment;
- any known instability, solubility or phase boundaries;
- the range supported by precedent or scouting data;
- the resolution at which the factor can be set and measured.
Avoid using the widest possible bounds by default. A model that spans ambient pressure to a reactor limit, for example, may spend much of the campaign resolving behaviour in regions the process team would never adopt. Equally, bounds copied from a literature procedure can exclude useful conditions for a different substrate or scale.
A boundary result is information. If the best observed condition sits at a permitted limit, do not automatically declare an optimum. Review whether the boundary is chemically or operationally fixed; if it is not, the next campaign may need an expanded range.
Match the factors to the development stage
Variable selection should evolve as the project matures.
| Stage | Primary question | Typical emphasis |
|---|---|---|
| Scouting | Which families are viable? | Broad categorical choices and a few informative numeric factors |
| Screening | Which factors and interactions matter? | Deliberate contrasts, controls and enough coverage to estimate effects |
| Optimisation | Where are the best feasible conditions? | Refined ranges around promising regions; relevant constraints |
| Robustness | How sensitive is the process near the chosen point? | Small deliberate perturbations, noise factors and replication |
| Scale-up | Will performance transfer? | Mixing, addition rate, heat transfer, concentration and hold times |
DoE guidance similarly begins with the objective, then the process variables and levels, before selecting a design. The NIST experimental-design handbook distinguishes comparative, screening, response-surface and other objectives because no single design answers all of them.
Work backwards from the budget
Every additional factor expands the space that the campaign must cover. This does not mean that high-dimensional campaigns are impossible, but it raises the burden on the design, model and interpretation.
Prioritise candidate factors using four questions:
- Plausibility: is there a chemical or process reason for an effect?
- Decision value: would learning its effect change the route or operating conditions?
- Controllability: can it be set reproducibly at the intended scale?
- Cost: what material, time, analytical or operational burden does it add?
When the list is still too long, split the work into stages. Use a screening design or focused scouting campaign to remove clearly inactive or impractical choices, then optimise the reduced space. Keep enough budget for controls, repeats, failed runs and confirmation experiments.
A worked example: coupling-condition development
Suppose a team has a workable coupling at 62% assay yield with variable conversion and a persistent dehalogenated impurity. The immediate decision is whether there is a catalyst system worth advancing.
A defensible first space might include:
- ligand identity as a categorical factor;
- base identity as a categorical factor;
- temperature over a range compatible with substrate stability;
- catalyst loading over a range the project could accept;
- assay yield and dehalogenated impurity as separate responses.
Reaction time could initially be fixed if all runs are sampled at a standard endpoint. Solvent might remain fixed if solubility scouting has already established a clear practical choice—or it might need inclusion if base behaviour changes markedly across solvents. Stirring rate should be recorded and controlled; it should become a factor only if heterogeneity or mass transfer is a credible concern.
The point is not that these are universally “the right variables”. The point is that each inclusion and exclusion follows from the current decision, chemistry and constraints.
Write a campaign brief before generating experiments
A compact brief makes assumptions visible to the whole team:
Decision: select a catalyst system for substrate scope work
Responses: assay yield (maximise), impurity A (minimise)
Factors: ligand, base, temperature, catalyst loading
Constraints: temperature ≤ 100 °C; approved solvent only
Fixed conditions: concentration, equivalents, reaction time
Context recorded: plate, position, reagent lots, operator, analysis batch
Budget: 24 development runs + 4 repeats/confirmations
Stop rule: predefined response targets or exhausted budget
Review this before the first run and whenever new evidence changes the space. Bayesian optimisation can adapt the next experiment to incoming results, as demonstrated in synthetic-reaction studies such as Shields and co-workers, but it cannot rescue ambiguous responses, infeasible categories or missing experimental context.
Before you start
- Can every factor be set and recorded reliably?
- Are units, category names and missing values defined?
- Are unsafe or incompatible combinations excluded?
- Does each range reflect real operational choices?
- Are important responses measured separately?
- Is there budget for repeats and confirmation?
- Will the resulting data answer the stated decision?
Good variable selection is not a one-off statistical exercise. It is the translation of chemical judgement into an experimental space that a team can execute, analyse and refine.
To see how those definitions move through the product, follow the AutoRxn walkthrough. You can also compare Bayesian optimisation and DoE, launch AutoRxn, or contact the team to discuss a campaign.
References and further reading
- NIST/SEMATECH, e-Handbook of Statistical Methods: Choosing an experimental design.
- Shields, B. J. et al., “Bayesian reaction optimization as a tool for chemical synthesis”, Nature 590, 89–96 (2021).