Robot learning · humanoids · physical AI
Safety validation has to move at the speed of the robot.
A new model, payload, site, task, or control strategy can change the assumptions behind yesterday's safety work. Saphira connects those changes to the hazards, requirements, tests, and evidence that need attention before the next deployment decision.
The validation problem
A safety case cannot be a snapshot.
Standards remain an essential foundation. But general-purpose robots operate across changing tasks and environments, so teams also need a traceable, evidence-backed argument for the system that exists now—not the one described months ago.
Model or software update
Which behavior assumptions, hazards, and scenarios need review?
New sensor or payload
Which mitigations and prior test results still apply?
New task or deployment site
How does the operating context change the safety argument?
Field event or intervention
Which scenario should be reproduced, expanded, and tested next?
A closed loop for validation
From real behavior to a reviewable deployment decision.
Use what happens in physical tests and remote-supervised operation to decide what to analyze, replay, test, and revalidate next.
- 01Evidence flow
Operate and test
Physical tests and remote-supervised operation produce telemetry, interventions, and observed behavior.
Inputs to Saphira · not a remote-operation product
- 02Evidence flow
Capture the event
A near miss, operator intervention, unexpected behavior, or failed test becomes a reviewable event.
Event · context · provenance
- 03Evidence flow
Map the impact
Saphira connects the event to affected assumptions, hazards, requirements, scenarios, and controls.
Traceability · change impact
- 04Evidence flow
Replay and verify
Teams turn the event into simulation scenarios and targeted physical tests with explicit thresholds.
Simulation · physical test
- 05Evidence flow
Update the safety case
Source-linked results return to the safety argument, with gaps surfaced for expert review.
Evidence · gaps · review
- 06Decision
Make the deployment decision
People decide whether the current system and operating conditions are supported by sufficient evidence.
Human release decision
The loop is an operating model. Remote-operation systems, test rigs, simulators, and deployed robots remain external evidence sources; Saphira connects their outputs to the safety model.
For robot learning teams
The policy improves. Its evaluation history should compound.
Saphira sits beside the learning and control stack. It connects policy and system changes to failure analysis, evaluation scenarios, thresholds, and evidence—then turns physical and remotely supervised behavior into the next regression suite.
Saphira does not train robot policies or control devices. It provides the evaluation and assurance layer that helps teams decide whether a changing system is ready for a specific task and deployment context.
Evidence in
- Policy, model, and system changes
- Task definitions and deployment assumptions
- Simulation rollouts and physical-test results
- Telemetry, failures, and operator interventions
Decisions out
- Structured failure modes and affected assumptions
- Scenario suites, thresholds, and coverage gaps
- Results linked to their source and system baseline
- Evidence for expert review and deployment decisions
01
Change
A new policy checkpoint, model, task, payload, control strategy, or deployment context.
02
Evaluate
Run targeted scenarios across simulation and physical test with explicit thresholds.
03
Deploy
Release against a reviewable evidence baseline and defined operating assumptions.
04
Observe
Capture field behavior, telemetry, failures, and remote-operation interventions.
05
Regress
Turn what happened into structured failure modes and repeatable evaluations.
06
Improve
Return targeted scenarios, data needs, and constraints to the learning and engineering loop.
Interface boundary
Device and agent interfaces make hardware operable. Learning systems improve behavior. Saphira addresses the adjacent lifecycle question: what could fail, what must be evaluated, what changed, and what evidence supports the next release?
Product evidence sequence
One event. One connected chain of evidence.
Follow one event from physical or remotely supervised operation through change-impact analysis, targeted simulation and test, and an updated evidence baseline.
Frame 01
Observed behavior
Event context, operator intervention, and physical-test evidence
Run 0421 · Station 04Remote supervisedPhysical test feedIntervention00:18.200:26.8Event
Unexpected reach near shared workspace
Operator action
Motion paused · manual retreat
Min dist.
0.42 m
Force
18 N
Frame 02
Change-impact view
Affected assumptions, hazards, requirements, and scenarios
Change impactAnalysis complete6 affectedSource event
EVT-421 · Reach intervention
Assumption
Human separation
Hazard
Unexpected contact
Requirement
Stop envelope
Scenario
Shared reach · person enters
Evidence gap
New checkpoint untested
Frame 03
Replay and targeted test
Scenario reproduction, thresholds, and result provenance
Scenario replay12 variantsTargeted suite
Baseline replayPassLate human entryPassOccluded approachFailPayload +4 kgQueuedSimulation · top view
S-08robotpersonSeparation threshold0.36 m · FAILFrame 04
Updated safety case
Source-linked evidence, remaining gaps, and expert review
Deployment baseline · v0.18Review requiredEvidence
31 / 34
Passed
11 / 12
Open gaps
3
EV-188Physical intervention traceLinkedSIM-642Occluded approach variants1 failedREQ-029Protective separation behaviorReviewRelease remains human-approved.
Open review
Go deeper
Published materials
Technical workflow
Virtual testing, SIL, and Task FMEA for humanoid robots
ReadCustomer case study
De-risking humanoid robotics at 1X
ReadCategory thesis
The risk engineering layer for autonomous machines
ReadEvidence architecture