The thesis

Physical AI needs a validation factory

Every serious robotics program eventually hits the same wall. The system works in the demo, and then somebody has to prove it is safe—not once, but every time the model, the payload, the site, or the control stack changes. That proof is currently produced by hand.

Validation is the bottleneck, not the model

The hard part of deploying a robot is no longer getting it to perform the task. It is establishing, to a standard an assessor and an insurer will accept, what the machine will do in the cases nobody demonstrated. That work is done today by small teams of safety engineers producing documents: a hazard analysis, a set of derived requirements, a test plan, a traceability matrix, a safety case. The documents are correct on the day they are signed and stale shortly after.

This is why adopting a new architecture is so expensive. The cost is not the integration itself; it is that everything downstream of the change has to be re-argued by people. In a field where nobody yet knows which compute platform, which middleware, or which policy architecture wins, a validation process that punishes change this severely is a tax on the entire industry.

A factory, not a document

The alternative is to treat validation the way we treat any other production problem: as a line with stations, inputs, outputs, and throughput. Ingest the system. Generate the hazard analysis from it. Derive the requirements. Generate and run the scenarios in simulation. Assemble the evidence. Deploy, and route what happens in the field back to the second station.

Once that loop exists, the interesting property is not that any single artifact is produced faster. It is that the marginal cost of re-validating approaches zero, which changes what engineering organizations can afford to do. You can change your sensor suite. You can ship a new policy checkpoint. You can qualify a new site. Each of those becomes a run of the line rather than a program of work.

How it works today

  • A hazard analysis written once, early, by whoever was available
  • Requirements that drift from the design within a quarter
  • Test plans authored by hand, covering what someone thought to test
  • A safety case reconstructed at the end from a year of scattered decisions
  • Any architecture change re-opens all of it

How a factory works

  • Hazard analysis generated from the system under test, and regenerated when it changes
  • Requirements derived from hazards, allocated and traceable both directions
  • Scenarios generated per requirement and executed in simulation
  • Evidence accumulated continuously, current as of the last run
  • An architecture change is a re-run, not a restart

Why simulation and safety have to be one system

Plenty of organizations can run a simulator, and plenty can write a safety case. The reason those two capabilities rarely sit in the same place is that they demand different expertise, and the teams that have one usually treat the other as a formality. Simulation without safety competence produces impressive coverage numbers against scenarios that were never tied to a hazard. Safety competence without simulation produces arguments that cannot be tested at the scale physical AI requires.

Saphira exists at that intersection deliberately. Our hazard analysis is generated from the system, including video from physical machines. Our scenario execution is orchestrated in the simulator the robotics industry actually uses. The same ROS 2 control interfaces we validate against in simulation are the ones running on the deployed machine. The loop closes because those pieces were built to be one system rather than integrated after the fact.

Where this goes

The physical AI stack is consolidating quickly on the compute and simulation layers, and it is entirely open on the layer that decides whether any of it can be deployed around people. That layer is validation. We think it becomes infrastructure—something every robotics company runs continuously rather than something a handful of them commission at the end of a program—and we are building it that way.

See the factory run on your system

Bring a robot, a vehicle, or a model checkpoint and we will show you what comes out the other end.

Stay updated with Saphira

Get the latest news and updates delivered to your inbox.