234 lines
12 KiB
Markdown
234 lines
12 KiB
Markdown
|
|
# Calibration Audit — 2026-07-27
|
||
|
|
|
||
|
|
## Scope and decision rule
|
||
|
|
|
||
|
|
This audit covers pre-prototype numerical and rigid-body simulation evidence
|
||
|
|
only. It does not contain physical prototype, independent F/T, fixture, or
|
||
|
|
human-subject results, and none of its values may be copied into the manuscript
|
||
|
|
Results section.
|
||
|
|
|
||
|
|
The calibration code was first committed as
|
||
|
|
`4503a12bf10886902e69d4be3874a76d3be577d6`. All three reported v2 raw batches
|
||
|
|
record that commit with `dirty=false`. The independent metric code and configs
|
||
|
|
were tested with 98 passing tests and one explicitly unsupported legacy MuJoCo
|
||
|
|
demo skipped.
|
||
|
|
|
||
|
|
The decision terms **verified contract**, **provisional numerical setting**,
|
||
|
|
**frozen setting**, and **not freeze-ready** follow
|
||
|
|
`docs/calibration/CALIBRATION_POLICY.md`.
|
||
|
|
|
||
|
|
The companion formula-linked audit workbook is
|
||
|
|
`outputs/calibration-20260727/calibration_audit_2026-07-27.xlsx`. It contains
|
||
|
|
13 visually inspected worksheets and has SHA-256
|
||
|
|
`24df38f67ba1aa428928348561b452db181f963ed30ae18a570aabc4bd8ad920`.
|
||
|
|
|
||
|
|
## Traceable batches
|
||
|
|
|
||
|
|
| Study | Completed | Plan hash | Metric-config hash | Row hash | Source-data SHA-256 |
|
||
|
|
|---|---:|---|---|---|---|
|
||
|
|
| H1 calibration v2 | 60/60 | `e86b4aec90912954fdee97834dee08dd630647a89000a24402fa56f25b3aea83` | `2be0e67feda3f862c3510b3318af47437977c2b87a17ab810ce68a49212d5215` | `9e274b01db38c97112ecd1e7b50050d5c342cf316a22ea9b2d2c0ea4acf347d4` | `85110ad917d6b933656571b15c764b26d6f8215bea6d6baa338d759ed51fc0ee` |
|
||
|
|
| H2 calibration v2 | 576/576 | `87f797edb5d2d6378a42b95616fca8ed733d0c81de5e22c422b006ead7071b12` | `caaf8c3b1c8180eb19afb98d986e8a54b8a897bd1ab706f78e39c41441932a31` | `fcac031716e967d3832d8eb6f770f580bd68bb6c8f715d55d04b5878f86f0fbf` | `e3a9bc29003193b559fbfc595ce2595b19a29a7135b9f14057128a8c5f504710` |
|
||
|
|
| Bilateral energy/gain v2 Stage A | 36/36 | `c10fa2cd84392f9dbf62809b0960bfb38d1e7aac4791ef0345a7353c45d8e062` | `37ca1e72ae55eb2a4594ef8e6395302d76ebfea87c4bca43615e86b3eeb89ec2` | `3f3f7f09387f02fe88c1b070815bd108f7e6bf63a091861ae62c1fac1dcdf421` | H3 `baa82701ea4c3e116657b9ef4d12aa795849957934f9dc5d58d5811405936da7`; H4 `47a9cc05f22724edf078a525e31ace6332b6048ab36ed72a005a50e8ecb79453`; diagnostics `a765cb6cc87c8fc89849c4d0a5d2560a768769169caa625baf41ddfd18f7af63` |
|
||
|
|
|
||
|
|
All batches validated with zero failed trials and no validator warnings.
|
||
|
|
|
||
|
|
## H1 — SEW retargeting
|
||
|
|
|
||
|
|
### Design corrections verified
|
||
|
|
|
||
|
|
- Five trajectory strata are now separate: nominal, wider valid reach,
|
||
|
|
valid near-singularity, intentional upper-reach clipping, and master
|
||
|
|
joint-limit stress.
|
||
|
|
- Each stratum has three genuinely different seeded trajectory instances.
|
||
|
|
All four methods receive a bit-identical master trajectory within a pair.
|
||
|
|
The audit found one trajectory hash per pair and 15 unique hashes across the
|
||
|
|
15 pairs.
|
||
|
|
- Reach clipping, active limits, geometry degeneracy, low manipulability, pose
|
||
|
|
failure, and differential invalidity are distinct machine-readable fields.
|
||
|
|
No invalid sample in v2 had an unexplained reason.
|
||
|
|
|
||
|
|
### Calibration outcome
|
||
|
|
|
||
|
|
| Method | Trials with \(C_r=1\) | Main observation |
|
||
|
|
|---|---:|---|
|
||
|
|
| SEW | 6/15 | Passed nominal, wider valid reach, and valid near-singularity; failed the intentionally invalid reach-clip and joint-limit strata |
|
||
|
|
| Task-priority IK | 6/15 | Same trajectory-level composite outcomes as SEW in every pair |
|
||
|
|
| Bounded DLS IK | 6/15 | Same trajectory-level composite outcomes as SEW in every pair |
|
||
|
|
| Scaled joint-space map | 15/15 | Mapping branch remained numerically valid in several strata, but task-space error exceeded the locked calibration tolerances |
|
||
|
|
|
||
|
|
The paired difference \(C_{r,\mathrm{SEW}}-C_{r,\mathrm{taskIK}}\) was zero in
|
||
|
|
all 15 pairs. Every method had \(D_r=0\), including all valid trajectories.
|
||
|
|
Consequently, this calibration does not support the H1 superiority hypothesis
|
||
|
|
and cannot calibrate or freeze the discontinuity threshold.
|
||
|
|
|
||
|
|
The valid near-singularity stratum is no longer mixed with reach clipping:
|
||
|
|
SEW, task-priority IK, and bounded DLS had 100% valid samples and zero reach
|
||
|
|
clips; the SEW slave-output low-manipulability fraction was approximately
|
||
|
|
0.502. The intentional upper-reach-clip stratum had a clip fraction of 1.0
|
||
|
|
and `reach_clipped_upper` as the primary invalid reason. The joint-limit
|
||
|
|
stratum had no reach clipping and correctly reported `joint_limit_active`.
|
||
|
|
|
||
|
|
Runtime tails over all v2 samples were:
|
||
|
|
|
||
|
|
| Method | P50 | P95 | P99 | Maximum |
|
||
|
|
|---|---:|---:|---:|---:|
|
||
|
|
| Scaled joint-space | 0.140 ms | 0.159 ms | 0.235 ms | 0.475 ms |
|
||
|
|
| Bounded DLS IK | 0.191 ms | 1.693 ms | 1.760 ms | 5.156 ms |
|
||
|
|
| SEW | 0.845 ms | 10.290 ms | 12.007 ms | 15.643 ms |
|
||
|
|
| Task-priority IK | 0.322 ms | 10.139 ms | 10.423 ms | 10.713 ms |
|
||
|
|
|
||
|
|
These are workstation numerical timings, not a hardware real-time claim.
|
||
|
|
|
||
|
|
### H1 decision
|
||
|
|
|
||
|
|
**Not freeze-ready.** Before a locked H1 study:
|
||
|
|
|
||
|
|
1. add valid branch-transition and near-degeneracy trajectories that can
|
||
|
|
actually exercise \(D_r\);
|
||
|
|
2. declare a target control period and separate target construction,
|
||
|
|
recovery, and differential-map timing;
|
||
|
|
3. tune both SEW and the primary task-priority baseline only on a new
|
||
|
|
calibration split;
|
||
|
|
4. require a nonzero number of eligible increments in every continuity
|
||
|
|
stratum; and
|
||
|
|
5. rerun with enough independent trajectories for a precision-based paired
|
||
|
|
confidence interval.
|
||
|
|
|
||
|
|
## H2 — residual-wrench inversion
|
||
|
|
|
||
|
|
### Pairing and conditioning audit
|
||
|
|
|
||
|
|
The v2 plan contains 16 physical-data groups. Each group has 36
|
||
|
|
\(\ell_c\times\lambda\times\)method candidates and one unique seed hash.
|
||
|
|
Independent array hashes confirmed that the wrench reference, joint velocity,
|
||
|
|
truth Jacobian, estimator Jacobian, and sensor noise are identical across all
|
||
|
|
36 candidates within each group.
|
||
|
|
|
||
|
|
Numerical rank and operational ill-conditioning are now separate. No sample was
|
||
|
|
numerically rank deficient, but the minimum scaled singular value reached
|
||
|
|
approximately \(9.83\times10^{-4}\), and 39.9% of samples at the provisional
|
||
|
|
candidate were below the preregistered operational threshold.
|
||
|
|
|
||
|
|
### Parameter scan
|
||
|
|
|
||
|
|
For scaled DLS, trial-mean force/moment RMSE at the strongest candidates was:
|
||
|
|
|
||
|
|
| \(\ell_c\) (m) | \(\lambda\) | Force mean / max (N) | Moment mean / max (Nm) |
|
||
|
|
|---:|---:|---:|---:|
|
||
|
|
| 0.2 | 0.02 | 3.051 / 9.336 | 0.722 / 1.478 |
|
||
|
|
| 0.2 | 0.03 | 2.848 / 8.229 | 0.704 / 1.269 |
|
||
|
|
| 0.2 | 0.05 | **2.763 / 7.322** | 0.705 / **1.126** |
|
||
|
|
| 0.3 | 0.03 | 3.033 / 8.058 | 0.772 / 1.283 |
|
||
|
|
| 0.4 | 0.02 | 3.184 / 8.448 | 0.812 / 1.362 |
|
||
|
|
|
||
|
|
At \(\ell_c=0.2\) m and \(\lambda=0.05\):
|
||
|
|
|
||
|
|
| Method | Force mean / max (N) | Moment mean / max (Nm) |
|
||
|
|
|---|---:|---:|
|
||
|
|
| Scaled DLS | 2.763 / 7.322 | 0.705 / 1.126 |
|
||
|
|
| Undamped SVD | 6.064 / 25.502 | 1.248 / 4.281 |
|
||
|
|
| No bias correction | 3.132 / 7.512 | 0.786 / 1.140 |
|
||
|
|
| No friction correction | 2.965 / 7.091 | 0.759 / 1.153 |
|
||
|
|
|
||
|
|
The paired DLS-minus-undamped mean differences were \(-3.301\) N and
|
||
|
|
\(-0.544\) Nm, but DLS was better in only 10 of 16 physical-data groups. It
|
||
|
|
was deliberately worse in several well-modelled, low-noise cells and much
|
||
|
|
better in the low-singular/model-error tail. This is a bias--variance result,
|
||
|
|
not a universal accuracy result. The largest undamped retained outlier was
|
||
|
|
25.502 N / 4.281 Nm; it was not removed.
|
||
|
|
|
||
|
|
### H2 decision
|
||
|
|
|
||
|
|
**Provisional numerical setting:** use scaled DLS with
|
||
|
|
\(\ell_c=0.2\) m and \(\lambda=0.05\) only for the next synthetic pilot.
|
||
|
|
Keep \(\ell_c=0.2,\lambda=0.03\) and undamped SVD as sensitivity conditions.
|
||
|
|
|
||
|
|
**Not a frozen physical calibration.** The length scale must ultimately be
|
||
|
|
anchored to robot geometry and physical calibration, and H2 cannot support a
|
||
|
|
paper wrench-accuracy claim before independent six-axis F/T truth, TCP/F/T
|
||
|
|
transforms, torque conversion, payload/friction calibration, timestamp
|
||
|
|
alignment, and causal acceleration estimation are available.
|
||
|
|
|
||
|
|
## H3/H4 — bilateral mapping and final-port energy supervision
|
||
|
|
|
||
|
|
### Contracts already verified
|
||
|
|
|
||
|
|
The v1 stored-array audit verified:
|
||
|
|
|
||
|
|
- fixed-branch \(A\)-port virtual-work error at numerical precision;
|
||
|
|
- same-input proposed-versus-matched-wrench behavior, with their contact
|
||
|
|
normalized-power-mismatch difference approximately zero;
|
||
|
|
- 96/96 tank trials passing the independent final-applied-port audit;
|
||
|
|
- maximum preclip floor deficit of zero;
|
||
|
|
- maximum accounting and software-preclip error of
|
||
|
|
\(6.94\times10^{-18}\) J; and
|
||
|
|
- zero downstream torque modification after the audited projection.
|
||
|
|
|
||
|
|
These are verified implementation contracts, not H3/H4 performance
|
||
|
|
acceptance.
|
||
|
|
|
||
|
|
The v1 free-space power normalizer was only about \(2\times10^{-4}\) J, so its
|
||
|
|
normalized mismatch was not suitable as a primary endpoint. Metric schema v2
|
||
|
|
therefore always reports absolute mismatch and marks normalized H3 values
|
||
|
|
invalid below a 1 mJ activity gate.
|
||
|
|
|
||
|
|
### Stage A outcome
|
||
|
|
|
||
|
|
Stage A scanned feedback strength 0.2/0.35/0.5 and initial headroom
|
||
|
|
0.005/0.020/0.040 J at two wall stiffness levels.
|
||
|
|
|
||
|
|
- 36/36 independent H4 audits passed.
|
||
|
|
- Maximum accounting and software-preclip error was
|
||
|
|
\(1.39\times10^{-17}\) J.
|
||
|
|
- Downstream modification and projected floor deficit were both zero.
|
||
|
|
- \(D_{\mathrm{proj}}\) remained 0.756--0.905, with a mean of 0.837.
|
||
|
|
- Projection intervention occupied 0.328--0.531 of samples.
|
||
|
|
- Every trial reached the configured 80 N contact-force limit.
|
||
|
|
- Slave tracking RMSE was 0.807--1.035 rad.
|
||
|
|
|
||
|
|
The audit therefore verifies the final-port implementation, while the tested
|
||
|
|
contact envelope is unsuitable for selecting a transparency setting. Choosing
|
||
|
|
the numerically smallest \(D_{\mathrm{proj}}\) cell would merely freeze a
|
||
|
|
force-limited, poorly tracked trajectory.
|
||
|
|
|
||
|
|
### H3/H4 decision
|
||
|
|
|
||
|
|
- **Verified contract:** differential-dual virtual work and final-applied-port
|
||
|
|
energy accounting.
|
||
|
|
- **Not freeze-ready:** feedback strength, energy bounds, H4 distortion
|
||
|
|
threshold, contact envelope, network envelope, and H3 superiority margin.
|
||
|
|
- **Stage B not executed:** its configuration explicitly requires an accepted
|
||
|
|
Stage A profile. Running it with the placeholder middle cell would violate
|
||
|
|
the calibration policy.
|
||
|
|
|
||
|
|
The next bilateral calibration must first establish a stable fixture envelope
|
||
|
|
without force limiting or large tracking error, then introduce a separate,
|
||
|
|
controlled energy-challenging excitation. Contact stability and energy-budget
|
||
|
|
excitation must not be forced by the same aggressive trajectory. Only after
|
||
|
|
that gate passes should the directional network screen cover forward/return
|
||
|
|
delay, jitter, and loss.
|
||
|
|
|
||
|
|
## Freeze record
|
||
|
|
|
||
|
|
| Item | Decision | Permitted use now |
|
||
|
|
|---|---|---|
|
||
|
|
| H1 thresholds and SEW superiority margin | Not freeze-ready | trajectory/debug development only |
|
||
|
|
| H2 scaled DLS \(\ell_c=0.2\) m, \(\lambda=0.05\) | Provisional | next synthetic pilot only |
|
||
|
|
| \(A^\top\) virtual-work implementation | Verified contract | deterministic regression gate |
|
||
|
|
| Final-output energy accounting | Verified contract | deterministic regression gate |
|
||
|
|
| H3 normalized endpoint | Definition repaired; margin not frozen | use gated value plus absolute mismatch in calibration |
|
||
|
|
| Feedback gain and tank energy window | Not freeze-ready | continue calibration |
|
||
|
|
| Bilateral Stage B network matrix | Blocked by Stage A selection gate | do not execute as a locked or confirmatory study |
|
||
|
|
|
||
|
|
## Required next execution order
|
||
|
|
|
||
|
|
1. Redesign and rerun H1 continuity trajectories.
|
||
|
|
2. Run the H2 provisional setting on a larger held-out synthetic pilot while
|
||
|
|
preserving the 16-group pairing contract.
|
||
|
|
3. Establish a stable, unsaturated fixture/contact envelope.
|
||
|
|
4. Design a distinct bounded energy-challenging excitation and rerun Stage A.
|
||
|
|
5. Replace Stage B placeholder haptic values only after Stage A passes.
|
||
|
|
6. Build the prototype measurement chain and complete physical calibration.
|
||
|
|
7. Freeze new `locked` JSON plans, metric configs, margins, exclusions, and a
|
||
|
|
clean source commit before collecting confirmatory data.
|