Purpose and scope
Measure how real NVIDIA GR00T action-chunk inference behaves under synchronous, asynchronous prefetch, and deadline-aware control. A six-joint SO100 arm grasps, lifts, transports, releases, and settles a cube in a target region in MuJoCo. Physics continues while inference is pending. This is simulation evidence; it does not establish physical-robot safety or sim-to-real performance.
Training uses contact-physics expert demonstrations with independent seed splits. Learned-policy results below are separated from expert data generation and transport test fixtures.
Execution environment
| Component | Observed value |
|---|---|
| GPU | NVIDIA A40 |
| VRAM GiB | 44.431640625 |
| CUDA | 12.8 |
| PyTorch | 2.9.0+cu128 |
| Python | 3.12.3 (main, Aug 14 2025, 17:47:21) [GCC 13.3.0] |
| Preflight passed | True |
| numpy | 1.26.4 |
| mujoco | 3.3.7 |
| scipy | 1.15.3 |
| pyzmq | 27.0.1 |
| msgpack | 1.1.0 |
| pyarrow | 19.0.1 |
| imageio | 2.37.0 |
| imageio-ffmpeg | 0.6.0 |
| flash_attention | 2.8.3 |
| torchcodec | 0.8.0 |
| driver | name, driver_version, memory.total [MiB] NVIDIA A40, 570.195.03, 46068 MiB |
Resource measurements
| Metric | Observed |
|---|---|
| Peak VRAM GiB | 37.921 |
| Minimum free disk GiB | 4.774 |
| Peak power W | 258.4 |
| Peak GPU temperature C | 61.0 |
Training
Loss indicates optimization progress, not task success. Checkpoint selection must use validation evidence; final acceptance uses held-out seeds.
Failure analysis and corrective work
Earlier candidates were retained as negative evidence. Their failures were used to change the data, control path, and acceptance process before the final model was evaluated.
| Run | Observed failure |
|---|---|
| gpu-smoke-v2/20260929T145149Z-2ea73da9 | placed but did not release; stable 30Hz control |
| gpu-horizon-validation/20260929T145456Z-87138805 | two timeouts and a rejected 0.061892-radian joint-limit overshoot |
| gpu-fullchunk-validation/20260929T150153Z-35af9a32 | three of three true task successes, but roughly 15 percent intentional inference holds fail the deployment fallback gate |
| gpu-fullchunk-deadline/20260929T150357Z-afa3d8c0 | timeout and rejected 0.074142-radian overshoot |
| 20260929T160522Z-190d441b | Five episodes exceeded the 1% predicted joint-value projection fraction |
| 20260929T163323Z-bc4aa72d | Observed wrist projection fraction 1.53846%; interrupted to reserve corrective training time |
Corrections implemented
- Demonstration jaw target moved from mechanical bound -0.174 to 0.0 radians
- Removed artificial waypoint dwell
- Retained full release and retreat in every demonstration
- Policy preparation now happens before timed control on socket owner thread
- Errors identify joint and action index
- Optional native GR00T RTC prefix conditioning added; actual inference validation pending
- Lossless image compression with total decompression budget
Independent checks completed
- 53 local tests and 53 GPU tests pass
- V1, V2, and eight-step V2 checkpoints fully verified locally
- Final credential-free evidence archive contains 332 entries and is preserved with a SHA-256 checksum
- Offline processor restoration verified from 4.3MB credential-free bundle
- Failure and success videos saved with physics replay checks
- V3 passed 100-trial held-out production acceptance with 92 successes and zero joint-limit corrections
- Two real workstation-to-GPU campaigns completed 5 of 5 task successes each through an encrypted SSH tunnel
Research and implementation sources
V1, V2, and V3 candidate comparison
| Candidate | Training change | Validation | Acceptance trials | Success | Safety projection | Release status | Research finding |
|---|---|---|---|---|---|---|---|
| V1 | Initial contact-physics demonstrations | 3 of 3 successes at full chunks, but about 15% fallback; shorter-horizon trials timed out or exceeded the safety allowance | Not admitted | Validation-only evidence | Rejected overshoots of 0.061892 and 0.074142 rad | Not approved | The task could complete, but the demonstrations and scheduling contract were not safe enough for acceptance |
| V2 | Safer jaw target, removed dwell, complete release and retreat | 10 of 10 candidate validation successes | 100 | 98 of 100 (98.0%) | Five trials exceeded the 1% per-episode projection gate | Not approved | High task success did not compensate for repeated wrist-limit corrections |
| V3 | Interior expert, 0.08 rad waypoint margin, warm start from V2 | 10 of 10 successes with zero corrections | 100 | 92 of 100 (92.0%), 95% CI 85.0 to 95.9% | Zero corrections across all 100 trials | Approved | Lower headline success than V2, but the complete safety, timing, provenance, and reliability contract passed |
Closed-loop run
Evidence: v1-runs\20260929T145456Z-87138805\episodes.jsonl. Sweep completed: False.
Checkpoint identity: 293b8cac12ae896728fc2285d6d7825e314f60facb97553498ef18befe6cb259. Prefetch lead: 0.25 seconds.
| Strategy | K | Added delay s | Trials | Success % | 95% CI % | RPC p50 s | RPC p95 s | RPC p99 s | Fallback fraction | Mean jerk rad/s^3 | Mean Hz | Deadline misses | Successes/clock hour |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deadline | 24 | 0 | 1 | 0.0 | 0.0 to 79.3 | 0.190 | 0.201 | 0.206 | 0.018 | 102.331 | 30.001 | 0 | 0.0 |
| deadline | 32 | 0 | 2 | 0.0 | 0.0 to 65.8 | 0.193 | 0.208 | 0.222 | 0.040 | 138.025 | 30.003 | 0 | 0.0 |
Closed-loop run
Evidence: v2-runs\20260929T160522Z-190d441b\episodes.jsonl. Sweep completed: True.
Checkpoint identity: 10a84840ba5f28b4588bab2b2eb95ba5e4c6e9325554dd596b91844dabe79030. Prefetch lead: 0.23 seconds.
| Strategy | K | Added delay s | Trials | Success % | 95% CI % | RPC p50 s | RPC p95 s | RPC p99 s | Fallback fraction | Mean jerk rad/s^3 | Mean Hz | Deadline misses | Successes/clock hour |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deadline | 40 | 0 | 100 | 98.0 | 93.0 to 99.4 | 0.199 | 0.206 | 0.219 | 0.016 | 66.783 | 30.002 | 6 | 286.3 |
Closed-loop run
Evidence: v3-runs\20260929T172231Z-f7aee121\episodes.jsonl. Sweep completed: True.
Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 0.23 seconds.
| Strategy | K | Added delay s | Trials | Success % | 95% CI % | RPC p50 s | RPC p95 s | RPC p99 s | Fallback fraction | Mean jerk rad/s^3 | Mean Hz | Deadline misses | Successes/clock hour |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deadline | 40 | 0 | 100 | 92.0 | 85.0 to 95.9 | 0.200 | 0.208 | 0.217 | 0.016 | 58.865 | 30.003 | 9 | 261.6 |
Closed-loop run
Evidence: v3-validation\20260929T171645Z-d9da7813\episodes.jsonl. Sweep completed: True.
Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 0.23 seconds.
| Strategy | K | Added delay s | Trials | Success % | 95% CI % | RPC p50 s | RPC p95 s | RPC p99 s | Fallback fraction | Mean jerk rad/s^3 | Mean Hz | Deadline misses | Successes/clock hour |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deadline | 40 | 0 | 10 | 100.0 | 72.2 to 100.0 | 0.197 | 0.205 | 0.234 | 0.017 | 55.451 | 30.003 | 1 | 290.7 |
Closed-loop run
Evidence: wan\20260929T175311Z-16639156\episodes.jsonl. Sweep completed: True.
Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 0.75 seconds.
| Strategy | K | Added delay s | Trials | Success % | 95% CI % | RPC p50 s | RPC p95 s | RPC p99 s | Fallback fraction | Mean jerk rad/s^3 | Mean Hz | Deadline misses | Successes/clock hour |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deadline | 40 | 0 | 5 | 100.0 | 56.6 to 100.0 | 0.609 | 0.948 | 1.053 | 0.201 | 110.366 | 30.078 | 763 | 277.1 |
Closed-loop run
Evidence: wan\20260929T175609Z-50ee183c\episodes.jsonl. Sweep completed: True.
Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 1.15 seconds.
| Strategy | K | Added delay s | Trials | Success % | 95% CI % | RPC p50 s | RPC p95 s | RPC p99 s | Fallback fraction | Mean jerk rad/s^3 | Mean Hz | Deadline misses | Successes/clock hour |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| deadline | 40 | 0 | 5 | 100.0 | 56.6 to 100.0 | 0.609 | 0.809 | 0.941 | 0.164 | 141.762 | 30.105 | 742 | 288.1 |
Production acceptance
v2-runs\20260929T160522Z-190d441b\acceptance.json : FAIL
{
"min_episodes": 30,
"min_success": 0.8,
"max_fallback": 0.05,
"max_deadline_fraction": 0.05,
"min_hz": 28.5,
"max_projection_fraction": 0.01
}- {"K": 40, "added_delay_s": 0, "backend": "real", "strategy": "deadline"}: excessive joint-limit projection
v3-runs\20260929T172231Z-f7aee121\acceptance.json : PASS
{
"min_episodes": 30,
"min_success": 0.8,
"max_fallback": 0.05,
"max_deadline_fraction": 0.05,
"min_hz": 28.5,
"max_projection_fraction": 0.01
}Visual evidence
Recorded motion and outcomes
dd848d18426c4da6ad5023f4893cdfe0.2177b2e7ea974837bd66c324d39e6eca.0fc8a8adb66a491998872ed27a1f28b9.012c3063df0141adbb8a5307d8ade018.c1ec4db161c34cd9852da0a8cad34786.b243a9b8cdda44309ba9852223e94101.6f5c1142ab0f423ba3660e90caed21e1.Run notes and remaining limitations
- V1 proved that the learned policy could complete the simulated task, but its action bounds and fallback behavior were not acceptable for release.
- V2 raised held-out task success to 98 percent, yet five trials exceeded the one percent per-episode joint-limit projection gate. The release was correctly rejected.
- V3 traded some headline task success for a safer learned distribution. It passed every declared gate with 92 successes in 100 fresh trials and zero joint-limit corrections.
- The real workstation-to-GPU route completed 10 of 10 tasks across baseline and tuned runs. Its 0.81 to 0.95 second p95 RPC latency left too many fallback ticks, so that Internet path is task-functional but not timing-qualified.
- The production approval applies to the declared MuJoCo SO100 task and checkpoint. Physical robot safety and sim-to-real performance remain separate validation work.
Interpretation
- A single simulated task with narrow initial-position randomization does not demonstrate general robotics capability.
- Command jerk and measured joint jerk are distinct. The reported jerk is computed from measured positions within each episode using actual physics timestamps.
- A 30-trial empirical 80% success threshold does not imply a 95% lower confidence bound of 80%. Confidence intervals are reported separately.
- CPU, rendering, RPC, and scheduling overhead are included in closed-loop timing. Offline visualization is excluded from timed control.