Robot learning · humanoids · physical AI

Safety validation has to move at the speed of the robot.

A new model, payload, site, task, or control strategy can change the assumptions behind yesterday's safety work. Saphira connects those changes to the hazards, requirements, tests, and evidence that need attention before the next deployment decision.

The validation problem

A safety case cannot be a snapshot.

Standards remain an essential foundation. But general-purpose robots operate across changing tasks and environments, so teams also need a traceable, evidence-backed argument for the system that exists now—not the one described months ago.

Model or software update

Which behavior assumptions, hazards, and scenarios need review?

New sensor or payload

Which mitigations and prior test results still apply?

New task or deployment site

How does the operating context change the safety argument?

Field event or intervention

Which scenario should be reproduced, expanded, and tested next?

A closed loop for validation

From real behavior to a reviewable deployment decision.

Use what happens in physical tests and remote-supervised operation to decide what to analyze, replay, test, and revalidate next.

  1. 01Evidence flow

    Operate and test

    Physical tests and remote-supervised operation produce telemetry, interventions, and observed behavior.

    Inputs to Saphira · not a remote-operation product

  2. 02Evidence flow

    Capture the event

    A near miss, operator intervention, unexpected behavior, or failed test becomes a reviewable event.

    Event · context · provenance

  3. 03Evidence flow

    Map the impact

    Saphira connects the event to affected assumptions, hazards, requirements, scenarios, and controls.

    Traceability · change impact

  4. 04Evidence flow

    Replay and verify

    Teams turn the event into simulation scenarios and targeted physical tests with explicit thresholds.

    Simulation · physical test

  5. 05Evidence flow

    Update the safety case

    Source-linked results return to the safety argument, with gaps surfaced for expert review.

    Evidence · gaps · review

  6. 06Decision

    Make the deployment decision

    People decide whether the current system and operating conditions are supported by sufficient evidence.

    Human release decision

The loop is an operating model. Remote-operation systems, test rigs, simulators, and deployed robots remain external evidence sources; Saphira connects their outputs to the safety model.

For robot learning teams

The policy improves. Its evaluation history should compound.

Saphira sits beside the learning and control stack. It connects policy and system changes to failure analysis, evaluation scenarios, thresholds, and evidence—then turns physical and remotely supervised behavior into the next regression suite.

Saphira does not train robot policies or control devices. It provides the evaluation and assurance layer that helps teams decide whether a changing system is ready for a specific task and deployment context.

Evidence in

  • Policy, model, and system changes
  • Task definitions and deployment assumptions
  • Simulation rollouts and physical-test results
  • Telemetry, failures, and operator interventions

Decisions out

  • Structured failure modes and affected assumptions
  • Scenario suites, thresholds, and coverage gaps
  • Results linked to their source and system baseline
  • Evidence for expert review and deployment decisions
  1. 01

    Change

    A new policy checkpoint, model, task, payload, control strategy, or deployment context.

  2. 02

    Evaluate

    Run targeted scenarios across simulation and physical test with explicit thresholds.

  3. 03

    Deploy

    Release against a reviewable evidence baseline and defined operating assumptions.

  4. 04

    Observe

    Capture field behavior, telemetry, failures, and remote-operation interventions.

  5. 05

    Regress

    Turn what happened into structured failure modes and repeatable evaluations.

  6. 06

    Improve

    Return targeted scenarios, data needs, and constraints to the learning and engineering loop.

Interface boundary

Device and agent interfaces make hardware operable. Learning systems improve behavior. Saphira addresses the adjacent lifecycle question: what could fail, what must be evaluated, what changed, and what evidence supports the next release?

Product evidence sequence

One event. One connected chain of evidence.

Follow one event from physical or remotely supervised operation through change-impact analysis, targeted simulation and test, and an updated evidence baseline.

  1. Frame 01

    Observed behavior

    Event context, operator intervention, and physical-test evidence

    Run 0421 · Station 04
    Remote supervised
    Physical test feedIntervention
    00:18.200:26.8

    Event

    Unexpected reach near shared workspace

    Operator action

    Motion paused · manual retreat

    Min dist.

    0.42 m

    Force

    18 N

  2. Frame 02

    Change-impact view

    Affected assumptions, hazards, requirements, and scenarios

    Change impact
    Analysis complete

    Source event

    EVT-421 · Reach intervention

    6 affected

    Assumption

    Human separation

    Hazard

    Unexpected contact

    Requirement

    Stop envelope

    Scenario

    Shared reach · person enters

    Evidence gap

    New checkpoint untested

  3. Frame 03

    Replay and targeted test

    Scenario reproduction, thresholds, and result provenance

    Scenario replay
    12 variants

    Targeted suite

    Baseline replayPass
    Late human entryPass
    Occluded approachFail
    Payload +4 kgQueued

    Simulation · top view

    S-08
    robotperson
    Separation threshold0.36 m · FAIL
  4. Frame 04

    Updated safety case

    Source-linked evidence, remaining gaps, and expert review

    Deployment baseline · v0.18
    Review required

    Evidence

    31 / 34

    Passed

    11 / 12

    Open gaps

    3

    EV-188Physical intervention traceLinked
    SIM-642Occluded approach variants1 failed
    REQ-029Protective separation behaviorReview

    Release remains human-approved.

    Open review