AI is turning laboratories into closed-loop discovery systems
Models can now propose hypotheses, select experiments, control instruments and learn from results. Scientific value still depends on calibration, safety, sample identity, uncertainty, provenance and independent replication.
What is happening?
A closed-loop laboratory does more than automate a fixed recipe. It uses each result to decide what experiment should happen next. An AI system may search the literature, suggest a hypothesis, choose conditions, send a validated protocol to robots, analyze the measurement and update its model. This can reduce repetitive work and explore complex spaces more intelligently. It can also optimize noise, misidentify a sample or repeat a calibration error at machine speed. The scientific standard therefore cannot be “the robot completed the run.” The result must remain safe, traceable, uncertainty-aware and reproducible by an independent method or laboratory.
Why this trend is moving
- 01AI co-scientist systems can generate hypotheses, research plans and experimentally testable proposals from literature and data.
- 02Robotic laboratories can prepare samples, operate instruments, analyze measurements and select the next experiment without waiting for a human between every cycle.
- 03Active learning and Bayesian optimization make expensive experiments more selective rather than merely faster.
- 04Open orchestration platforms and modular hardware are lowering the cost of connecting instruments into closed loops.
- 05National laboratories and public programs are funding autonomous experimentation for materials, catalysts, biology and energy research.
- 06Safety, calibration, sample provenance, benchmark quality and cross-laboratory reproducibility are becoming the main barriers to trustworthy scale.
What this means in practice
- The complete laboratory loop—not the language model or robot alone—is the system that must be evaluated.
- Natural-language proposals should pass through typed protocols, unit checks and deterministic device constraints before physical execution.
- Sample identity, reagent lots, calibration, raw measurements and every data transformation need machine-actionable provenance.
- Active-learning objectives must represent the real scientific question rather than a convenient proxy such as yield alone.
- Independent safety interlocks should constrain materials, quantities, temperature, pressure, motion and waste handling.
- Negative, aborted and failed experiments should remain in the evidence base because they define operating boundaries.
- Promising findings need held-back confirmation and independent replication before they become broad scientific claims.
What the headline leaves out
This is the practical technical view: how the system is put together, where it can fail, and what a real deployment asks from the team running it.
How it is built
A production autonomous-laboratory system begins with a bounded scientific objective, falsification criteria and an operational design domain for materials, devices and hazards. Samples, reagents, instruments, calibration states and software receive persistent identities. Literature and prior experimental data support hypothesis generation, while active learning or Bayesian optimization selects candidate experiments with explicit uncertainty. A deterministic compiler converts the selected plan into validated quantities, units, device commands and controls. Robots execute the protocol inside independent safety limits, instruments preserve raw measurements and analysis pipelines estimate quality and uncertainty. The loop decides whether to repeat, explore, recalibrate, ask a scientist or stop. Findings are then confirmed with held-back materials, alternative methods or another laboratory, with full provenance preserved.
How inference behaves
Closed-loop science connects four functions: scientific reasoning, experiment selection, physical execution and evidence update. Language models or domain models can generate hypotheses and plans. Active-learning algorithms choose experiments expected to reduce uncertainty or improve a declared objective. Laboratory orchestration software schedules robots and instruments, while deterministic controllers handle units, timing and device constraints. Analysis converts raw signals into measurements with uncertainty. Those results update the model and influence the next choice. Each transition can introduce error, so sample tracking, calibration, controls, versioning and post-run validation are part of the scientific mechanism rather than administrative additions.
What the tests can miss
A credible evaluation measures validated information gain, intervention-free campaign completion, protocol accuracy, sample-identity integrity, repeatability, uncertainty calibration, optimization efficiency, accessible parameter-space coverage, safety compliance, drift detection, failure classification, replication, provenance completeness and complete cost. Comparisons should declare the reference strategy—random search, grid search, human expertise or conventional design of experiments. Throughput and a discovered optimum are insufficient when measurement noise, proxy choice, unreported human rescue or missing failed runs could explain the result.
What deployment involves
Start with an AI research copilot or human-approved experiment-selection loop. Move to bounded autonomy only after protocols, devices, sample tracking, calibration and recovery are qualified inside a narrow operational design domain. Use deterministic execution layers and independent safety interlocks. Preserve human review for novel hazards, ambiguous scientific goals and campaigns that cross qualified boundaries. Promote findings only after held-back confirmation and independent replication. Scale through modular instrument interfaces and portable provenance records rather than one-off scripts that work only in a single laboratory.
Where the risks sit
Autonomous laboratories connect AI systems to physical equipment, inventories, research data, remote facilities and sometimes hazardous materials. Access to instrument commands, sample databases and protocol libraries should follow least privilege. Model-generated text, retrieved papers and external data must not directly authorize physical execution. Signed software, protected calibration, network segmentation, controlled remote access and immutable audit records reduce manipulation risk. Research data may contain commercial secrets, unpublished discoveries, genomic information or export-controlled material and requires appropriate access and retention controls.
What it really costs
The cost of a closed-loop laboratory includes robots, instruments, integration, safety engineering, sample tracking, consumables, cleaning, calibration, compute, expert review, failed experiments, maintenance, downtime and independent replication. Faster cycle time can still produce poor economics when analytical equipment is the bottleneck or technicians spend substantial time recovering the platform. The useful denominator is validated information gained or independently confirmed findings—not experiment count, robot uptime or model tokens.
What the evidence supports
The trend is credible and increasingly broad. Google’s Co-Scientist has generated biomedical hypotheses that were reviewed and experimentally tested. NIST operates autonomous formulation and materials programs while developing standards for samples, instruments, data and model integration. Recent peer-reviewed systems report closed-loop reaction optimization, biotechnology, perovskite fabrication, polymer discovery, X-ray alignment and experiment-theory interaction. A-Lab, CAMEO, Polybot, RoboChem-Flex and AMASE demonstrate different combinations of robotics, active learning and scientific models. The boundaries are equally clear: successful systems remain bespoke and scientifically narrow; recent safety and benchmarking work emphasizes operational limits, intervention, precision, drift, provenance and replication. The evidence supports bounded autonomous campaigns—not unsupervised general scientific authority.
How it works in practice
AI can help laboratories propose hypotheses, choose experiments, operate instruments and learn from results, but an autonomous loop is only scientifically useful when sample identity, calibration, uncertainty, safety, provenance and independent replication remain stronger than the optimization objective. The goal is not maximum experiment count. It is trustworthy information gained per constrained unit of time, material and risk.
How the parts work together
The headline technology is only one part of the product. Reliability, security and cost are usually decided by the handoffs around it.
- 01
Scientific objective and validity boundary
Define the research question, measurable target, admissible evidence, excluded conditions, stopping rules and what would falsify the working hypothesis. Optimization without a scientific boundary can efficiently find an artefact.
- 02
Operational design domain and safety envelope
Declare allowed materials, concentrations, temperatures, pressures, instruments, waste paths and human-supervision requirements. Independent interlocks constrain physical execution before any AI-selected experiment reaches hardware.
- 03
Sample, instrument and protocol identity
Assign persistent identifiers to samples, reagents, instruments, calibration states, software, methods and containers. Every measurement must remain linked to what was actually prepared and observed.
- 04
Hypothesis and experiment selection
Literature, prior data and domain models generate candidate hypotheses or conditions. Active learning, Bayesian optimization or agent planning selects the next experiment with explicit uncertainty and exploration constraints.
- 05
Deterministic protocol compilation
A validated execution layer converts the scientific plan into instrument-specific commands, units, volumes, timing and sequencing. It rejects impossible, unsafe or ambiguous instructions rather than improvising at the device boundary.
- 06
Physical execution and measurement
Robots and instruments prepare, manipulate and characterize samples while monitoring run state, calibration, environmental conditions and exceptions. Raw measurements are preserved before downstream transformation.
- 07
Analysis, uncertainty and loop decision
The system checks data quality, propagates uncertainty, compares alternatives and decides whether to repeat, explore, exploit, recalibrate, ask a scientist or stop. Failed and negative experiments remain part of the evidence.
- 08
Replication, provenance and controlled learning
Promising results are reproduced with held-back materials, independent methods or another laboratory. The complete evidence package records lineage, interventions and changes before results are promoted into models, publications or product decisions.
Estimate the limits before the demo
These equations are planning tools rather than substitutes for testing. They help expose a design that is unlikely to fit its hardware, budget, reliability or risk limits.
The useful denominator is information gain, not experiment count
research efficiency = validated information gain ÷ (time + material + energy + human intervention + risk cost) Throughput can rise while scientific value falls if the loop repeats correlated measurements, chases instrument drift or optimizes a weak proxy. Information gain should be tied to uncertainty reduction or a declared decision.
- Ten diverse experiments can be more useful than one hundred near-duplicates.
- A calibration run has value when it prevents false discovery.
- A negative result can eliminate a large region of hypothesis space.
Measurement uncertainty limits optimization confidence
observed variance ≈ process variance + measurement variance + preparation variance An optimizer can exploit noise when measurement and preparation uncertainty are not estimated. Replicates, controls and calibration are required before small differences are treated as better chemistry or biology.
- A one-percent yield improvement is meaningless when assay variation is three percent.
- Batch-to-batch reagent variation can look like a model discovery.
- Instrument drift can pull an active learner toward the wrong region.
Closed loops compound hidden failure probabilities
valid campaign probability ≈ Π(valid stageᵢ) Planning, sample preparation, transfer, measurement, analysis and data linkage can each appear reliable in isolation. A small undetected error at every stage compounds across long campaigns, especially when the resulting data trains the next decision.
- A mislabeled vial can contaminate every later model update.
- One unit-conversion error can create an apparently coherent optimization path.
- Independent checkpoints reduce correlated propagation.
A capable model does not make an autonomous laboratory
A self-driving laboratory connects algorithms to physical experiments. The system may search literature, propose hypotheses, select conditions, compile protocols, move samples, control instruments, analyze measurements and choose the next experiment. Recent work spans materials synthesis, chemistry, biotechnology, formulation, microscopy and synchrotron beamlines.
The impressive part is often the model or robot demonstration. The scientific result, however, depends on every connection between them. A wrong unit, stale calibration, swapped sample, partial cleaning cycle or silent instrument warning can become training data for the next decision. The loop can then become confidently self-reinforcing around an artefact.
Evaluate the campaign as one versioned scientific system. Hardware, software, methods, reagents, environment, optimization policy and human interventions all belong in the evidence record.
- Name the complete laboratory configuration behind every claim.
- Separate hypothesis quality from execution reliability.
- Preserve raw data and instrument state.
- Treat every automated transformation as part of the scientific method.
Language-model proposals require a scientific compiler
AI co-scientist systems can synthesize literature and propose hypotheses or protocols. That is different from generating a hardware-ready experiment. Natural language leaves units, tolerances, reagent identity, sequence, controls, contamination risk and instrument capability underspecified.
A reliable architecture separates scientific intent from deterministic execution. The proposal is converted into a typed protocol with validated quantities, allowed devices, preconditions, expected observations and failure handling. Domain tools should calculate stoichiometry, concentrations and instrument settings rather than relying on free-form arithmetic.
The compiler must be allowed to reject the plan. Asking a human is a successful control when required information is missing; speculative completion is not autonomy.
- Use structured protocol schemas with units and tolerances.
- Validate every command against device and material limits.
- Require positive and negative controls where the claim needs them.
- Keep model text away from direct actuator authority.
Active learning can optimize the wrong proxy with great efficiency
Bayesian optimization and active learning are valuable because experiments are expensive and parameter spaces are large. They can balance exploration of uncertain regions with exploitation of promising ones. The algorithm still depends on the objective and constraints chosen by people.
A single measured proxy may not represent the scientific or product goal. High yield can hide instability, toxicity, difficult purification or poor scale-up. Multi-objective optimization makes trade-offs visible, but only for objectives that were measured and correctly weighted.
The campaign should include uncertainty, calibration checks, replicate policy, novelty checks and stopping rules. An optimizer should not continue merely because budget remains.
- Declare objective functions before the campaign.
- Include feasibility and safety constraints.
- Measure uncertainty and replicate near decision boundaries.
- Review whether the proxy still represents the real goal.
Machine-actionable provenance is part of experimental validity
Autonomous laboratories generate data quickly, but speed magnifies the cost of weak metadata. A result should identify the sample, precursor lots, preparation steps, container, instrument, firmware, calibration, environmental conditions, raw signal, processing code and every later transformation.
NIST’s autonomous-laboratory and research-data work emphasizes interoperable sample, instrument, data and model standards. FAIR principles make data findable, accessible, interoperable and reusable, but reuse also requires enough context to judge whether two experiments are comparable.
Negative, aborted and failed runs should not disappear. They reveal operating boundaries, prevent repeated waste and help distinguish algorithmic failure from laboratory failure.
- Assign persistent sample and run identifiers.
- Record reagent lots, calibration and environmental state.
- Link processed values back to immutable raw measurements.
- Preserve failures and human interventions with reason codes.
A syntactically valid protocol can still be chemically or biologically unsafe
Laboratory autonomy introduces hazards that text-only agents do not face: pressure, heat, incompatible chemicals, aerosols, contamination, moving equipment, radiation, biological material and waste. A model may propose a plausible sequence without understanding the complete physical consequence.
Safety should be enforced through an operational design domain, device interlocks, inventory controls, ventilation and containment rules, maximum quantities, validated waste handling and emergency-stop logic. These controls should remain independent of the planning model.
Human review should scale with hazard and novelty. Running a known formulation within qualified bounds is different from synthesizing an unfamiliar compound or modifying a biological organism.
- Define permitted materials, quantities and operating ranges.
- Keep interlocks and emergency stops independent of AI.
- Block unapproved substitutions and unknown waste paths.
- Require hazard review when the scientific domain expands.
Calibration drift can masquerade as discovery
Closed-loop campaigns assume that a measurement today is comparable with one made earlier. Instruments drift, probes foul, liquids evaporate, robots wear, lighting changes and reagents age. An optimizer can learn these changes as if they were material properties.
Reference samples, blanks, control charts, scheduled calibration and drift detectors must be part of the loop. The system should know when uncertainty has grown enough to pause and recalibrate. Retrospective correction should remain traceable rather than silently rewriting the record.
Cross-instrument and cross-laboratory replication are especially important for claims that will leave the original platform.
- Interleave reference and blank measurements.
- Monitor drift against declared control limits.
- Re-run boundary results after calibration changes.
- Separate biological or chemical variation from instrument variation.
A discovered optimum must survive replication and alternative explanations
Self-driving-lab performance is often reported through throughput, optimization rate or the number of experiments saved. Those metrics are useful, but they do not establish novelty, mechanism, robustness or transfer.
A credible result reports accessible parameter space, degree of autonomy, intervention, precision, operational lifetime, material use and the comparison strategy. It also tests whether the finding persists across new reagent lots, randomization, held-back conditions, independent analysis and, where practical, another instrument or laboratory.
The scientific question may require causal evidence rather than a high-performing recipe. The autonomous system should distinguish optimization from explanation.
- Predefine success, novelty and replication criteria.
- Report intervention and failed-run rates.
- Use held-back materials and randomized execution order.
- Seek independent replication before broad claims.
Integration and exception handling determine whether autonomy saves time
A closed-loop laboratory requires instrument adapters, robotics, scheduling, sample storage, cleaning, calibration, safety engineering, data infrastructure and staff who can diagnose failures. Shared analytical equipment and slow characterization can become the true bottleneck even when planning is instant.
Low-cost modular platforms and open orchestration systems are reducing entry barriers, but interoperability remains incomplete. Vendor lock-in, unsupported drivers and bespoke protocol code can make the system difficult to transfer or maintain.
The business and scientific case should measure validated learning per complete campaign cost. Include setup, failed runs, consumables, instrument time, expert review, maintenance, downtime and independent replication.
- Automate the bottleneck, not merely the visible manual step.
- Prefer modular interfaces and portable data schemas.
- Track intervention and mean time to recovery.
- Compare with human-guided design of experiments and conventional automation.
What a benchmark worth believing should report
A performance number means little unless the workload, system configuration and quality bar are fixed. This is the minimum record a team should keep.
| Metric | How to measure it | Why it matters |
|---|---|---|
| Validated information gain | Quantify uncertainty reduction or decision value after replication rather than counting completed runs. | Experiment volume can rise without producing reliable knowledge. |
| Intervention-free campaign completion | Count campaigns completed without unplanned manual rescue, relabeling or protocol repair. | Hidden expert labor can dominate apparent autonomy. |
| Protocol compilation accuracy | Compare generated quantities, units, sequence, controls and device commands with qualified reference protocols. | A plausible plan can fail at the hardware boundary. |
| Sample-identity integrity | Audit chain of identity from precursor and container through measurement and processed result. | A mislabeled sample invalidates downstream learning. |
| Measurement repeatability | Run replicates, blanks and references across time and operators. | Optimization cannot distinguish signal from noise without repeatability. |
| Uncertainty calibration | Compare predicted intervals with observed outcomes across the parameter space. | Experiment selection depends on trustworthy uncertainty. |
| Optimization efficiency | Compare experiments and time required to reach a declared target against random, grid, expert or design-of-experiments baselines. | Acceleration needs a meaningful reference strategy. |
| Accessible parameter-space coverage | Report the safe and physically reachable region rather than the theoretical search space. | Claims can be inflated by counting conditions the platform cannot execute. |
| Safety-envelope compliance | Count rejected plans, interlock events, near misses and excursions outside qualified limits. | Scientific speed cannot be separated from physical safety. |
| Calibration and drift performance | Track reference-sample error, drift detection and recovery after recalibration. | Drift can become a false optimization signal. |
| Failure classification quality | Label preparation, device, measurement, analysis, data-linkage and scientific failures separately. | The next decision depends on understanding why a run failed. |
| Independent replication rate | Measure how many promoted findings survive held-back materials, alternative methods or another laboratory. | Closed-loop success inside one platform is not general scientific validity. |
| Provenance completeness | Audit required sample, instrument, protocol, software, calibration and transformation metadata. | Results without lineage cannot be reproduced or challenged. |
| Complete cost per validated finding | Include hardware, integration, consumables, instrument time, compute, intervention, maintenance, failed runs and replication. | Per-experiment cost understates discovery economics. |
Four sensible deployment patterns
AI research copilot
- Where it fits
- Literature review, hypothesis generation, protocol drafting and data analysis with scientists retaining execution authority.
- What you take on
- Fastest and lowest physical risk, but manual translation and experiment selection remain bottlenecks.
Human-approved closed loop
- Where it fits
- Automated preparation and measurement where the system recommends each next experiment for scientist approval.
- What you take on
- Strong control and learning visibility, with slower cycle time and review workload.
Bounded autonomous campaign
- Where it fits
- Qualified materials, instruments and parameter ranges with deterministic safety and automatic experiment selection.
- What you take on
- High throughput inside a narrow operational design domain; expansion requires requalification.
Federated autonomous laboratory network
- Where it fits
- Multiple specialized labs, simulations and agents coordinating through shared standards and evidence packages.
- What you take on
- Broadest capability and strongest replication opportunity, with major interoperability, identity, IP and governance complexity.
Where projects usually go wrong
The optimizer exploits measurement noise
What you see: The apparent optimum disappears under replication.
What to do: Estimate variance, use replicates and require improvement beyond measurement uncertainty.
A sample loses its identity
What you see: The result cannot be confidently linked to precursor lot, container or preparation path.
What to do: Use persistent identifiers, barcode or machine-readable transfers and chain-of-custody checks.
Natural-language protocol reaches hardware directly
What you see: Ambiguous units or sequence produce an unsafe or invalid run.
What to do: Compile into a typed validated protocol before execution.
Instrument drift becomes a discovery
What you see: Later samples appear systematically better as calibration moves.
What to do: Interleave references, monitor control charts and pause on drift.
The objective is a weak proxy
What you see: The platform improves yield while stability, toxicity or manufacturability worsens.
What to do: Use multi-objective constraints and review proxy validity.
Failed runs disappear from the dataset
What you see: The model repeatedly revisits unsafe or infeasible regions.
What to do: Preserve failed and aborted experiments with structured reason codes.
The model proposes an unsafe but syntactically valid experiment
What you see: Commands remain within API format but violate chemical or biological safety.
What to do: Enforce independent operational boundaries, inventory rules and interlocks.
Training data and validation become entangled
What you see: The loop tunes against the same cases later used to claim discovery.
What to do: Reserve held-back materials, conditions and replication paths.
The system confuses optimization with explanation
What you see: A high-performing recipe is presented as a causal mechanism.
What to do: Require mechanism-specific experiments and alternative-hypothesis testing.
A software update changes physical behavior silently
What you see: The same protocol produces a different sequence or instrument setting.
What to do: Pin versions, replay dry runs and requalify material changes.
Interoperability depends on fragile custom scripts
What you see: A device replacement or transfer to another lab breaks the campaign.
What to do: Use modular interfaces, declared schemas and portable provenance records.
Throughput hides complete campaign cost
What you see: The robot runs continuously while experts spend substantial time cleaning, recovering and validating.
What to do: Measure intervention, downtime, maintenance and replication in the denominator.
A checklist you can actually use
- Define the scientific claim, falsification condition and stopping rule.
- Declare the operational design domain and hazard class.
- List permitted materials, quantities, devices and environmental ranges.
- Keep emergency stops and safety interlocks independent of AI.
- Assign persistent identifiers to samples, reagents and runs.
- Record instrument, firmware, calibration and environment state.
- Preserve raw measurements before processing.
- Convert model proposals into typed validated protocols.
- Use domain tools for units, stoichiometry and device constraints.
- Declare objective functions and proxy limitations.
- Estimate preparation and measurement variance.
- Define replicate, blank and reference-sample policy.
- Track uncertainty and calibration of the experiment-selection model.
- Preserve negative, aborted and failed experiments.
- Separate scientific failure from hardware and data failure.
- Reserve held-back materials and conditions for confirmation.
- Randomize execution where order effects are possible.
- Require human review when hazard, novelty or uncertainty exceeds bounds.
- Test recovery from device, transfer and analysis failures.
- Version every model, protocol, driver and analysis pipeline.
- Require independent replication for promoted discoveries.
- Audit provenance completeness before publication or transfer.
- Calculate complete cost per validated finding.
Terms worth knowing
- Self-driving laboratory
- A laboratory that uses automation and algorithms to select, execute and analyze experiments in a feedback loop.
- Closed loop
- A process in which experimental results directly inform the selection of the next experiment.
- Active learning
- A method that selects new observations expected to improve a model most efficiently.
- Bayesian optimization
- Sequential optimization using a predictive surrogate and uncertainty-aware acquisition strategy.
- Operational design domain
- The materials, instruments, conditions and hazards within which autonomous operation is qualified.
- Protocol compiler
- A deterministic layer that converts a scientific plan into validated instrument-ready commands.
- Sample provenance
- The documented identity and history of a sample from inputs through preparation, measurement and transformation.
- Measurement traceability
- An unbroken documented chain connecting a result to references, calibrations and stated uncertainty.
- Acquisition function
- A rule used by an optimizer to choose the next experiment from predicted value and uncertainty.
- Exploration
- Selecting uncertain conditions to learn more about the experimental space.
- Exploitation
- Selecting conditions expected to perform well under the current model.
- Sim-to-real gap
- The difference between predicted or simulated behavior and physical experimental behavior.
- Negative result
- A valid experiment that does not support the target effect and still contributes scientific evidence.
- Independent replication
- Repetition using held-back materials, another method, instrument or laboratory to test whether a finding persists.
- Validated information gain
- A reduction in scientific uncertainty that survives quality checks and replication.
Primary references and technical starting points
These sources support the architecture, runtime, benchmark and security claims. Vendor capabilities can change, so the article records the distinction between established evidence, measured product behavior and editorial interpretation.
- 01 Nature: Accelerating scientific discovery with Co-Scientistnature.com
- 02 Google Research: AI co-scientistresearch.google
- 03 Nature Methods: What’s your hypothesis?nature.com
- 04 Scientific Reports: AutoLabs autonomous chemical experimentationnature.com
- 05 Communications Materials: Multi-agent autonomous materials labsnature.com
- 06 Nature Synthesis: RoboChem-Flexnature.com
- 07 Nature Communications: IvoryOS laboratory orchestrationnature.com
- 08 Nature: Autonomous closed-loop perovskite solar cellsnature.com
- 09 Nature Machine Intelligence: Agentic X-ray scientistnature.com
- 10 npj Robotics: Robotics in self-driving labsnature.com
- 11 Nature Chemical Engineering: Adaptive autonomous polymer discoverynature.com
- 12 Nature Communications: LLM agents for atomic force microscopynature.com
- 13 Scientific Reports: Autonomous biotechnology laboratorynature.com
- 14 Nature Communications: Performance metrics for self-driving labsnature.com
- 15 Nature Reviews Chemistry: Safe self-driving laboratoriesnature.com
- 16 Chemical Reviews: Self-driving laboratories for chemistry and materials sciencepubs.acs.org
- 17 Nature: A-Lab autonomous inorganic materials synthesisnature.com
- 18 Nature Communications: CAMEO closed-loop materials discoverynature.com
- 19 Science Advances: AMASE experiment-theory closed loopscience.org
- 20 NIST: Autonomous laboratoriesnist.gov
- 21 NIST: Autonomous Formulation Labnist.gov
- 22 NIST: Standards for modular autonomous laboratoriesnist.gov
- 23 NIST: Materials data and protocolsnist.gov
- 24 NIST Research Data Frameworknvlpubs.nist.gov
- 25 NIST: Data and specimen provenancenist.gov
- 26 ARPA-E: CATALCHEM-E autonomous laboratory programarpa-e.energy.gov
- 27 Argonne National Laboratory: Polybotcnm.anl.gov
- 28 Berkeley Lab: FORUM-AIforum-ai.lbl.gov
- 29 Berkeley Lab: Genesis Mission autonomous labslbl.gov
- 30 Berkeley Lab: A-Lab overviewnewscenter.lbl.gov
- 31 Berkeley Lab: OPAL autonomous biological discoverynewscenter.lbl.gov
- 32 Argonne: Real-time AI imaging for autonomous discoveryaps.anl.gov
- 33 Scientific Data: FAIR Guiding Principlesnature.com
- 34 arXiv: Safe-SDL safety boundariesarxiv.org
- 35 arXiv: Autonomous Laboratory Agent with deterministic executionarxiv.org
- 36 arXiv: Socratic agents for autonomous scientific discoveryarxiv.org
- 37 arXiv: Agent Laboratory research assistantsarxiv.org
- 38 arXiv: Hypothesis hunting with autonomous scientific agentsarxiv.org
- 39 arXiv: Benchmarking self-driving labsarxiv.org
- 40 arXiv: Human-guided autonomous materials explorationarxiv.org
- 41 Nature Chemical Engineering: Roadmap for autonomous nanomaterials experimentationnature.com
- 42 Berkeley Lab: AI assistant for energy materials discoverynewscenter.lbl.gov
- 43 NIST AI Risk Management Frameworknist.gov
- 44 NIST AI 600-1 Generative AI Profilenist.gov