Purpose and scope
Measure how real NVIDIA GR00T action-chunk inference behaves under synchronous, asynchronous prefetch, and deadline-aware control. A six-joint SO100 arm grasps, lifts, transports, releases, and settles a cube in a target region in MuJoCo. Physics continues while inference is pending. This is simulation evidence; it does not establish physical-robot safety or sim-to-real performance.
Training uses contact-physics expert demonstrations with independent seed splits. Learned-policy results below are separated from expert data generation and transport test fixtures.
Training
Loss indicates optimization progress, not task success. Checkpoint selection must use validation evidence; final acceptance uses held-out seeds.
Failure analysis and corrective work
Earlier candidates were retained as negative evidence. Their failures were used to change the data, control path, and acceptance process before the final model was evaluated.
| Run | Observed failure |
|---|---|
| gpu-smoke-v2/20260929T145149Z-2ea73da9 | placed but did not release; stable 30Hz control |
| gpu-horizon-validation/20260929T145456Z-87138805 | two timeouts and a rejected 0.061892-radian joint-limit overshoot |
| gpu-fullchunk-validation/20260929T150153Z-35af9a32 | three of three true task successes, but roughly 15 percent intentional inference holds fail the deployment fallback gate |
| gpu-fullchunk-deadline/20260929T150357Z-afa3d8c0 | timeout and rejected 0.074142-radian overshoot |
| 20260929T160522Z-190d441b | Five episodes exceeded the 1% predicted joint-value projection fraction |
| 20260929T163323Z-bc4aa72d | Observed wrist projection fraction 1.53846%; interrupted to reserve corrective training time |
Corrections implemented
- Demonstration jaw target moved from mechanical bound -0.174 to 0.0 radians
- Removed artificial waypoint dwell
- Retained full release and retreat in every demonstration
- Policy preparation now happens before timed control on socket owner thread
- Errors identify joint and action index
- Optional native GR00T RTC prefix conditioning added; actual inference validation pending
- Lossless image compression with total decompression budget
Independent checks completed
- 53 local tests and 53 GPU tests pass
- V1, V2, and eight-step V2 checkpoints fully verified locally
- Final credential-free evidence archive contains 332 entries and is preserved with a SHA-256 checksum
- Offline processor restoration verified from 4.3MB credential-free bundle
- Failure and success videos saved with physics replay checks
- V3 passed 100-trial held-out production acceptance with 92 successes and zero joint-limit corrections
- Two real workstation-to-GPU campaigns completed 5 of 5 task successes each through an encrypted SSH tunnel
Research and implementation sources
V1, V2, and V3 candidate comparison
| Candidate | Training change | Validation | Acceptance trials | Success | Safety projection | Release status | Research finding |
|---|---|---|---|---|---|---|---|
| V1 | Initial contact-physics demonstrations | 3 of 3 successes at full chunks, but about 15% fallback; shorter-horizon trials timed out or exceeded the safety allowance | Not admitted | Validation-only evidence | Rejected overshoots of 0.061892 and 0.074142 rad | Not approved | The task could complete, but the demonstrations and scheduling contract were not safe enough for acceptance |
| V2 | Safer jaw target, removed dwell, complete release and retreat | 10 of 10 candidate validation successes | 100 | 98 of 100 (98.0%) | Five trials exceeded the 1% per-episode projection gate | Not approved | High task success did not compensate for repeated wrist-limit corrections |
| V3 | Interior expert, 0.08 rad waypoint margin, warm start from V2 | 10 of 10 successes with zero corrections | 100 | 92 of 100 (92.0%), 95% CI 85.0 to 95.9% | Zero corrections across all 100 trials | Approved | Lower headline success than V2, but the complete safety, timing, provenance, and reliability contract passed |
Closed-loop run
Evidence: runs\20260929T160522Z-190d441b\episodes.jsonl. Sweep completed: True.
Checkpoint identity: 10a84840ba5f28b4588bab2b2eb95ba5e4c6e9325554dd596b91844dabe79030. Prefetch lead: 0.23 seconds.
| Strategy | K | Added delay s | Trials | Success % | 95% CI % | RPC p50 s | RPC p95 s | RPC p99 s | Fallback fraction | Mean jerk rad/s^3 | Mean Hz | Deadline misses | Successes/clock hour |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deadline | 40 | 0 | 100 | 98.0 | 93.0 to 99.4 | 0.199 | 0.206 | 0.219 | 0.016 | 66.783 | 30.002 | 6 | 286.3 |
Production acceptance
runs\20260929T160522Z-190d441b\acceptance.json : FAIL
{
"min_episodes": 30,
"min_success": 0.8,
"max_fallback": 0.05,
"max_deadline_fraction": 0.05,
"min_hz": 28.5,
"max_projection_fraction": 0.01
}- {"K": 40, "added_delay_s": 0, "backend": "real", "strategy": "deadline"}: excessive joint-limit projection
Visual evidence
Recorded motion and outcomes
0fc8a8adb66a491998872ed27a1f28b9.012c3063df0141adbb8a5307d8ade018.c1ec4db161c34cd9852da0a8cad34786.Run notes and remaining limitations
- V2 achieved 98 successes in 100 fresh held-out trials, with two step-limit outcomes.
- Five episodes exceeded the one percent per-episode predicted-value projection gate, so V2 was not released despite its high task success.
- The wrist-correction video shows why a safety gate must remain independent from aggregate task success.
Interpretation
- A single simulated task with narrow initial-position randomization does not demonstrate general robotics capability.
- Command jerk and measured joint jerk are distinct. The reported jerk is computed from measured positions within each episode using actual physics timestamps.
- A 30-trial empirical 80% success threshold does not imply a 95% lower confidence bound of 80%. Confidence intervals are reported separately.
- CPU, rendering, RPC, and scheduling overhead are included in closed-loop timing. Offline visualization is excluded from timed control.