EVALUATION NOTES / ROBOT LEARNINGMJAK Labs

GR00T cloud robotics: V1, V2, and V3

Researcher - Akhilesh Chandra

Date - 29 September 2026

ACCEPTANCE PASSEDRelease status
223Recorded closed-loop attempts
210Successful real-policy trials
A40 / H=40GPU / prediction horizon

Purpose and scope

Measure how real NVIDIA GR00T action-chunk inference behaves under synchronous, asynchronous prefetch, and deadline-aware control. A six-joint SO100 arm grasps, lifts, transports, releases, and settles a cube in a target region in MuJoCo. Physics continues while inference is pending. This is simulation evidence; it does not establish physical-robot safety or sim-to-real performance.

Training uses contact-physics expert demonstrations with independent seed splits. Learned-policy results below are separated from expert data generation and transport test fixtures.

Execution environment

ComponentObserved value
GPUNVIDIA A40
VRAM GiB44.431640625
CUDA12.8
PyTorch2.9.0+cu128
Python3.12.3 (main, Aug 14 2025, 17:47:21) [GCC 13.3.0]
Preflight passedTrue
numpy1.26.4
mujoco3.3.7
scipy1.15.3
pyzmq27.0.1
msgpack1.1.0
pyarrow19.0.1
imageio2.37.0
imageio-ffmpeg0.6.0
flash_attention2.8.3
torchcodec0.8.0
drivername, driver_version, memory.total [MiB] NVIDIA A40, 570.195.03, 46068 MiB

Resource measurements

MetricObserved
Peak VRAM GiB37.921
Minimum free disk GiB4.774
Peak power W258.4
Peak GPU temperature C61.0
Ten-second telemetry samples; short peaks between samples may be missed. Setup, training and evaluation are all included.
Ten-second telemetry samples; short peaks between samples may be missed. Setup, training and evaluation are all included.

Training

Loss indicates optimization progress, not task success. Checkpoint selection must use validation evidence; final acceptance uses held-out seeds.

Training losses read directly from saved trainer state.
Training losses read directly from saved trainer state.

Failure analysis and corrective work

Earlier candidates were retained as negative evidence. Their failures were used to change the data, control path, and acceptance process before the final model was evaluated.

RunObserved failure
gpu-smoke-v2/20260929T145149Z-2ea73da9placed but did not release; stable 30Hz control
gpu-horizon-validation/20260929T145456Z-87138805two timeouts and a rejected 0.061892-radian joint-limit overshoot
gpu-fullchunk-validation/20260929T150153Z-35af9a32three of three true task successes, but roughly 15 percent intentional inference holds fail the deployment fallback gate
gpu-fullchunk-deadline/20260929T150357Z-afa3d8c0timeout and rejected 0.074142-radian overshoot
20260929T160522Z-190d441bFive episodes exceeded the 1% predicted joint-value projection fraction
20260929T163323Z-bc4aa72dObserved wrist projection fraction 1.53846%; interrupted to reserve corrective training time

Corrections implemented

Independent checks completed

Research and implementation sources

V1, V2, and V3 candidate comparison

CandidateTraining changeValidationAcceptance trialsSuccessSafety projectionRelease statusResearch finding
V1Initial contact-physics demonstrations3 of 3 successes at full chunks, but about 15% fallback; shorter-horizon trials timed out or exceeded the safety allowanceNot admittedValidation-only evidenceRejected overshoots of 0.061892 and 0.074142 radNot approvedThe task could complete, but the demonstrations and scheduling contract were not safe enough for acceptance
V2Safer jaw target, removed dwell, complete release and retreat10 of 10 candidate validation successes10098 of 100 (98.0%)Five trials exceeded the 1% per-episode projection gateNot approvedHigh task success did not compensate for repeated wrist-limit corrections
V3Interior expert, 0.08 rad waypoint margin, warm start from V210 of 10 successes with zero corrections10092 of 100 (92.0%), 95% CI 85.0 to 95.9%Zero corrections across all 100 trialsApprovedLower headline success than V2, but the complete safety, timing, provenance, and reliability contract passed
Held-out closed-loop success by candidate. The dotted line is the 80 percent empirical release threshold.
Held-out closed-loop success by candidate. The dotted line is the 80 percent empirical release threshold.

Closed-loop run

Evidence: v1-runs\20260929T145456Z-87138805\episodes.jsonl. Sweep completed: False.

Checkpoint identity: 293b8cac12ae896728fc2285d6d7825e314f60facb97553498ef18befe6cb259. Prefetch lead: 0.25 seconds.

StrategyKAdded delay sTrialsSuccess %95% CI %RPC p50 sRPC p95 sRPC p99 sFallback fractionMean jerk rad/s^3Mean HzDeadline missesSuccesses/clock hour
deadline24010.00.0 to 79.30.1900.2010.2060.018102.33130.00100.0
deadline32020.00.0 to 65.80.1930.2080.2220.040138.02530.00300.0
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.

Closed-loop run

Evidence: v2-runs\20260929T160522Z-190d441b\episodes.jsonl. Sweep completed: True.

Checkpoint identity: 10a84840ba5f28b4588bab2b2eb95ba5e4c6e9325554dd596b91844dabe79030. Prefetch lead: 0.23 seconds.

StrategyKAdded delay sTrialsSuccess %95% CI %RPC p50 sRPC p95 sRPC p99 sFallback fractionMean jerk rad/s^3Mean HzDeadline missesSuccesses/clock hour
deadline40010098.093.0 to 99.40.1990.2060.2190.01666.78330.0026286.3
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.

Closed-loop run

Evidence: v3-runs\20260929T172231Z-f7aee121\episodes.jsonl. Sweep completed: True.

Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 0.23 seconds.

StrategyKAdded delay sTrialsSuccess %95% CI %RPC p50 sRPC p95 sRPC p99 sFallback fractionMean jerk rad/s^3Mean HzDeadline missesSuccesses/clock hour
deadline40010092.085.0 to 95.90.2000.2080.2170.01658.86530.0039261.6
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.

Closed-loop run

Evidence: v3-validation\20260929T171645Z-d9da7813\episodes.jsonl. Sweep completed: True.

Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 0.23 seconds.

StrategyKAdded delay sTrialsSuccess %95% CI %RPC p50 sRPC p95 sRPC p99 sFallback fractionMean jerk rad/s^3Mean HzDeadline missesSuccesses/clock hour
deadline40010100.072.2 to 100.00.1970.2050.2340.01755.45130.0031290.7
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.

Closed-loop run

Evidence: wan\20260929T175311Z-16639156\episodes.jsonl. Sweep completed: True.

Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 0.75 seconds.

StrategyKAdded delay sTrialsSuccess %95% CI %RPC p50 sRPC p95 sRPC p99 sFallback fractionMean jerk rad/s^3Mean HzDeadline missesSuccesses/clock hour
deadline4005100.056.6 to 100.00.6090.9481.0530.201110.36630.078763277.1
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.

Closed-loop run

Evidence: wan\20260929T175609Z-50ee183c\episodes.jsonl. Sweep completed: True.

Checkpoint identity: 9bb09bc4fe31f86c72eef8f7e4821a8a176f25454b2bad164e1f9c4d1dfe97e5. Prefetch lead: 1.15 seconds.

StrategyKAdded delay sTrialsSuccess %95% CI %RPC p50 sRPC p95 sRPC p99 sFallback fractionMean jerk rad/s^3Mean HzDeadline missesSuccesses/clock hour
deadline4005100.056.6 to 100.00.6090.8090.9410.164141.76230.105742288.1
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.
Real policy inference in wall-clock-paced simulation. Added completion delay is controlled injection, not a measurement of Internet latency.

Production acceptance

v2-runs\20260929T160522Z-190d441b\acceptance.json : FAIL

{
  "min_episodes": 30,
  "min_success": 0.8,
  "max_fallback": 0.05,
  "max_deadline_fraction": 0.05,
  "min_hz": 28.5,
  "max_projection_fraction": 0.01
}

v3-runs\20260929T172231Z-f7aee121\acceptance.json : PASS

{
  "min_episodes": 30,
  "min_success": 0.8,
  "max_fallback": 0.05,
  "max_deadline_fraction": 0.05,
  "min_hz": 28.5,
  "max_projection_fraction": 0.01
}

Visual evidence

fullchunk-success.png
media-v1\fullchunk-success.png. See accompanying metadata for expert versus learned trajectory provenance.
v1-rejected-command-hd.png
media-v1\v1-rejected-command-hd.png. See accompanying metadata for expert versus learned trajectory provenance.
v2-timeout-30094-hd.png
media-v2\v2-timeout-30094-hd.png. See accompanying metadata for expert versus learned trajectory provenance.
v2-validation-success-hd.png
media-v2\v2-validation-success-hd.png. See accompanying metadata for expert versus learned trajectory provenance.
v2-wrist-corrections-30061-hd.png
media-v2\v2-wrist-corrections-30061-hd.png. See accompanying metadata for expert versus learned trajectory provenance.
v3-acceptance-success-hd.png
media-v3\v3-acceptance-success-hd.png. See accompanying metadata for expert versus learned trajectory provenance.
v3-acceptance-timeout-hd.png
media-v3\v3-acceptance-timeout-hd.png. See accompanying metadata for expert versus learned trajectory provenance.

Recorded motion and outcomes

fullchunk-success: recorded trial. Exact logged-command reconstruction; not a second learned-policy evaluation. Episode: dd848d18426c4da6ad5023f4893cdfe0.
v1-rejected-command-hd: policy_error. Exact logged-command reconstruction; not a second learned-policy evaluation. Episode: 2177b2e7ea974837bd66c324d39e6eca.
v2-timeout-30094-hd: step_limit. Exact logged-command reconstruction; not a second learned-policy evaluation. Episode: 0fc8a8adb66a491998872ed27a1f28b9.
v2-validation-success-hd: success. Exact logged-command reconstruction; not a second learned-policy evaluation. Episode: 012c3063df0141adbb8a5307d8ade018.
v2-wrist-corrections-30061-hd: success. Exact logged-command reconstruction; not a second learned-policy evaluation. Episode: c1ec4db161c34cd9852da0a8cad34786.
v3-acceptance-success-hd: success. Exact logged-command reconstruction; not a second learned-policy evaluation. Episode: b243a9b8cdda44309ba9852223e94101.
v3-acceptance-timeout-hd: step_limit. Exact logged-command reconstruction; not a second learned-policy evaluation. Episode: 6f5c1142ab0f423ba3660e90caed21e1.

Run notes and remaining limitations

Interpretation