跳转至

Reconstruction 指标调研记录

本页保留详细公式、备选指标和文献来源,供需要追溯时查阅。日常实验请直接阅读 Reconstruction 评测指标

状态:讨论稿,不是 frozen evaluator contract。 本页只回答“每个 Recon 指标怎样算、适用于哪个 panel”;baseline 来源与接入状态见 Reconstruction baselines。论文唯一 confirmatory primary 暂定为 downstream Physics success,Recon metrics 支持 reconstruction claim。

先看结论

数据 我们实际比较什么 为什么
D3D-HOI · 正文主线 Orient. [°]、Loc. [cm]、Dim. [cm]、Rev. [°]、Pri. [cm] 这是 D3D-HOI native GT-CAD protocol;CAD 已给定,所以不评 Object CD / Moving-part CD
ARCTIC RGB · 可选 官方 hands+object 指标:hand MPJPE、两项 relative-root error、AAE、CDev、MDev、object vertex success 只在 ArcticNet-SF、HOPformer 和 Ours 能使用同一 view/split/evaluator 时整组加入
Generated video · supporting output coverage、self-consistency、failure 与 downstream Physics 没有冻结逐帧 GT 时不能声称 Recon accuracy

Contact distance 可以由几何计算;contact precision/recall/F1 只有在方法真的预测 contact set、并有独立 GT 时才 成立。下面保留完整公式和候选目录,日常执行只需看 Reconstruction 评测指标

1. 我们实际重建什么

目标输出为:

\[ \widehat{\mathcal R} =\{\widehat H_{1:T},\widehat T^o_{1:T},\widehat q_{1:T},\widehat s\}, \]

即 SMPL-X human、object world pose、articulation (q) 和 scale。Contact 分三类:

  • contact-intent input:GT、人工 review、VLM pseudo-label 或其他预处理标签;
  • achieved-contact diagnostic:由重建 mesh 几何导出的 gap/collision;
  • contact prediction:只有另建无人工 test intervention 的预测任务、独立 GT 和 P/R/F1 后才成立。

正式表逐行披露 contact provenance;由输入标签与重建几何计算的距离只称 label-qualified geometric gap

2. 三类入口与 GT / reference availability

同一数据源可有多个入口;表中协议是身份的一部分,不能只写“ARCTIC”。

Panel / protocol 输入 可用 GT 或 reference 不可假设存在 合法 lead metrics 主要 alignment
D3D-HOI Real RGB→Recon · P0 core 单目 RGB + GT CAD + GT pose/motion rendered object mask;所有 rows 权限相同 object base rotation/translation、dimensions、revolute/prismatic part motion、matched parametric CAD full-body GT、contact GT、独立 mesh geometry GT 原论文 native orientation/location/dimension/part-motion errors;coverage/failure 另作 ITT audit world/camera object frame;不作 PA
ARCTIC RGB→Recon · P1 time-gated 冻结 allo/exocentric 或 ego RGB protocol MoCap/MoSh++ SMPL-X full-body GT、双 MANO hands、object base + 1D articulation、contact pairs、calibrated cameras 不同 view/split 自由混用;把官方 MANO leaderboard 自动延伸为 full-body leaderboard official hands+object:\(\mathrm{MPJPE}_h\)\(\mathrm{MRRPE}_{r\to l}/\mathrm{MRRPE}_{r\to o}\)、AAE、object vertex [email protected]、CDev/MDev、\(\mathrm{ACC}_h/\mathrm{ACC}_o\);full-body supplement:world/camera + root-aligned MPJPE/PVE/root/temporal 只有完整 activation set 过 gate 才启用;run coverage/failure 不是 official metric
Generated Video→Recon 冻结 generator、prompt、seed、camera、asset 只有 generator 实际导出且冻结 的逐帧 human/object/q 才是 controlled reference;asset geometry/URDF 仅是资产 GT 仅凭 Infinigen/PartNet/GAPartNet/ArtVIP 资产自动得到 video-level human/contact/trajectory GT 有 reference 时用对应 accuracy;无 reference 时仅 consistency、coverage、downstream Physics frozen generator camera/world convention
ARCTIC captured/GT-like→Physics captured motion/reference,绕过 Recon controller reference Recon 输出/误差 Physics controller isolation canonical Studio frame

D3D-HOI 是 real-video articulated object benchmark; ARCTIC 的 consistent-motion 官方 leaderboard task 是 two hands + articulated object,但数据集本身同时提供 MoCap/MoSh++ SMPL-X full-body GT 和 calibrated cameras。ARCTIC 的手部 MPJPE_h 不得改名成 full-body MPJPE;full-body 是另建 supplement protocol。Ours 稳定输出是 SMPL-X,而官方 hands+object baselines 使用 MANO;在冻结并验证 MANO↔SMPL-X 21-joint 与 contact-vertex correspondence bridge 之前,Ours row 的 \(\mathrm{MPJPE}_h\)、CDev、MDev 与 vertex ACC 均为 Unavailable,不允许用 nearest vertex 事后选 mapping。Infinigen、PartNet-Mobility、GAPartNet 与 ArtVIP 是 asset sources,不是独立视频数据集; VideoGen 才是 generated-video pipeline。

3. 共同分母、对齐与统计

正式报告以 scheduled_manifest 为分母:crash、NaN、timeout、missing 和 early termination 均保留;continuous metric 只在有 GT 的真实 valid prefix 计算并同时报 coverage。每个预注册 case×seed 只有一个 canonical result;有多个 seed 时先在 case 内聚合,再按真实 CAD/object cluster 做 paired bootstrap。Revolute 与 prismatic 分开。当前跨数据集 manifest-driven ITT aggregate 尚未完成,已有数字只能标 pilot/historical。完整统计规则见 Evaluation strategy

4. Human metrics

\(J_{tj}\)\(V_{tv}\) 为 GT joint/vertex,帽号为预测。

\[ \operatorname{MPJPE}=\frac1{TJ}\sum_{t,j}\|\widehat J_{tj}-J_{tj}\|_2, \qquad \operatorname{PVE}=\frac1{TV}\sum_{t,v}\|\widehat V_{tv}-V_{tv}\|_2. \]

必须分别命名 world/camera、pelvis/root-aligned 和逐帧 similarity-Procrustes 后的 PA-MPJPE。 PA-MPJPE 只测 pose/shape diagnostic,因为它会移除 translation、rotation 和 scale。PVE 只在相同 topology 和 vertex correspondence 下成立。

root translation 与 orientation:

\[ e_{t}^{root}=\|\widehat t^{root}-t^{root}\|_2, \quad e_R^{root}=\cos^{-1}\!\left[ \operatorname{clip}\!\left(\frac{\operatorname{tr}(R^{\top}\widehat R)-1}{2},-1,1\right) \right]. \]

按 body/arm/left-hand/right-hand 分组报告 MPJPE/PVE,不能通过只报全身均值隐藏小尺度手部错误。 WHAM 提供 world-grounded human motion 中 MPJPE、PA-MPJPE、PVE、world trajectory、acceleration/jitter 和 foot sliding 的一手先例;ARCTIC 的官方 \(\mathrm{MPJPE}_h\) 是每只手去 root 后 21 joints 的 mm 误差,然后对两手汇总。

Metric 数据要求 优点 可投机/限制 位置
world/camera MPJPE 同 frame/world joint GT 直接支持 world-grounded claim camera/ground calibration 误差混入 full-body GT panel 的 main/secondary
root MPJPE root-aligned joint GT 隔离 pose 隐藏 global/root error ARCTIC official hands panel(仅激活时);其他 secondary
PA-MPJPE joint GT + similarity fit 隔离 shape/pose 移除论文关键 global error appendix diagnostic
PVE 同 topology mesh GT 表面密集 topology 不同不可比 secondary
root trans/orient global GT 支持 HOI layout 坐标约定敏感 key secondary
hand/body groups joint mapping 揭示局部退化 多重比较 appendix,预声明 groups
foot slide foot contact + FPS/ground 检查落地稳定性 contact detection 易循环定义 appendix/control quality

5. Object pose 与 geometry

\[ e_t=\|\widehat t-t\|_2, \qquad e_R=\cos^{-1}\!\left[ \operatorname{clip}\!\left(\frac{\operatorname{tr}(R^\top\widehat R)-1}{2},-1,1\right) \right]. \]

D3D-HOI 原论文报 object orientation 的 SO(3) relative angle(degree)、base-center location(cm)、dimension difference(cm),以及 revolute part-motion absolute error(degree)和 prismatic error(cm);这应是 D3D panel 的直接基线,而不是把 mask self-consistency 当 GT accuracy。 层级不能写错:official evaluator 对每个 video 计算一个 orientation/location/dimension error,只有 part motion 先在 video 内跨 frame 平均。原论文把 storage furniture 的 revolute/prismatic 分成两个 evaluation groups:pose/dimension 的 average 跨 9 个 groups;revolute motion 的 average 只跨 8 个 revolute groups、明确排除量纲为 cm 的 storage-prismatic;prismatic 列只汇总 17 个对应 videos,不能与 revolute degree 再平均。

对 video \(i\) 写为

\[ e_{\mathrm{dim},i}=\frac13\sum_{a\in\{w,h,d\}}|\widehat d_{ia}-d_{ia}|\;[\mathrm{cm}], \qquad e_{q,i}=\frac1{T_i}\sum_{t=1}^{T_i}|\widehat q_{it}-q_{it}|, \qquad \operatorname{unit}(e_{q,i})= \begin{cases} [\mathrm{deg}], & \text{revolute},\\ [\mathrm{cm}], & \text{prismatic}. \end{cases} \]

orientation 使用上式 \(e_{R,i}\),location 使用 \(e_{t,i}\)official evaluator 先对同 category 的 video 等权平均,原论文 Table 2 再给 category rows 与跨 category 的 macro average:

\[ \bar e_c=\frac1{|\mathcal I_c|}\sum_{i\in\mathcal I_c}e_i, \qquad \bar e_{\mathrm{macro}}=\frac1{|\mathcal C|}\sum_{c\in\mathcal C}\bar e_c. \]

正式 author-fidelity aggregate 必须复现这个 video→evaluation-group→macro 口径,并把 revolute/prismatic 分开;不能在看到结果后改 weighting。对 \(e_R,e_t,e_{\rm dim}\)\(\mathcal C\) 是 9 个 groups;对 revolute motion,\(\mathcal C\) 只含 8 个 revolute groups;prismatic motion 只在 17 个 storage-prismatic videos 内聚合。 该 reducer 已固化在 scripts/experiments/d3dhoi_native_recon_protocol.py;作者公开结果由 scripts/experiments/aggregate_d3dhoi_author_release.py 生成可追溯 aggregate,后续 Ours 的 256 个 result.pt 使用同一 reducer,禁止 degree/cm 混合平均。 本项目不在正文保留 automatic-CAD claim,因此不引入原论文 CAD-retrieval ablation,也不把 retrieved-correct subset 结果混入 GT-CAD main table。

3DADN 是 D3D panel 的 partial-output candidate,只比较 revolute state。3DHOI 论文在 239 个 revolute videos 上报告了它的结果,但该 adaptation 为满足 3DADN 的单调运动假设,使用 GT motion annotation 将输入切成 opening/closing 片段;论文还明确说明 3DADN 并非每帧都输出 motion state,reported error 只在 temporal optimization 有预测 motion parameters 的帧上计算。本地目前也没有 这批逐视频 predictions。因此正文必须先取得预测,再按 D3D native q-state 和统一 scheduled frames 重评;否则只能在附录标为 reported-only,并披露 GT-motion input privilege 及 video/frame coverage。 Orient./Loc./Dim./Pri. 不适用,不能补 GT 或补零。3DHOI 论文提出的 CubeOPT 另有 239 条我们的 author-code reproduction;它不是 3DADN,也不是作者发布结果。

5.1 3DHOI legacy nine-metric protocol(appendix only)

3DHOI paperofficial evaluator 在 239 个 revolute videos 上使用九项 error metrics。这是独立的 legacy panel,不与 D3D native GT-CAD 五项主表混合。对整体物体或可动部件的 10,000 个 surface samples \(P,Q\),official code 先将预测与 GT 各自归一化到相同最大直径,再作不估计尺度的 rigid ICP,最后计算 PyTorch3D 双向 squared Chamfer:

\[ \operatorname{CD}(P,Q)= \frac{1}{|P|}\sum_{p\in P}\min_{q\in Q}\|p-q\|_2^2+ \frac{1}{|Q|}\sum_{q\in Q}\min_{p\in P}\|q-p\|_2^2. \]

Object CD 与 Moving-part CD 因此是 normalized squared-distance score,不是 cm。若 ICP 给出 \((R_{\rm ICP},t_{\rm ICP})\),而预测/GT 未归一化前的最大直径为 \(\hat d,d\),则其 pose errors 为

\[ e_R=\cos^{-1}\!\left[\operatorname{clip}\!\left( \frac{\operatorname{tr}(R_{\rm ICP})-1}{2},-1,1\right)\right]\;[\mathrm{deg}],\quad e_t=\|t_{\rm ICP}\|_2\;[\text{normalized}],\quad e_s=1-\min\!\left(\frac{d}{\hat d},\frac{\hat d}{d}\right)\;[1]. \]

设 GT axis/origin 为 \((a,o)\),预测为 \((\hat a,\hat o)\),motion errors 为

\[ e_o=\frac{\| (\hat o-o)\times a\|_2}{\|a\|_2}\;[\text{normalized}],\quad e_a=\cos^{-1}\!\left[\operatorname{clip}\!\left( \frac{|a^\top\hat a|}{\|a\|_2\|\hat a\|_2},0,1\right)\right]\;[\mathrm{deg}],\quad e_{\rm dir}=\cos^{-1}\!\left[\operatorname{clip}\!\left( \frac{a^\top\hat a}{\|a\|_2\|\hat a\|_2},-1,1\right)\right]\;[\mathrm{deg}]. \]

State 是 degree absolute error;official code 对 CubeOPT/3DADN 先用一个 closed frame 重置零点,然后直接取 绝对差,没有 circular wrap,所以不能用另一个 wrapped-angle evaluator 补数。九项都是先逐帧、再逐 video 平均、最后跨 video 平均;3DADN State 的帧集例外必须另报 coverage。CubeOPT 还使用 GT object mask 和 GT moving-part mask,D3DHOI-GT-CAD 使用 GT object mask + GT CAD,因此该表是带 input-privilege 标记的 source-comparison,不是无条件 full-system ranking。

给定相同 model point set \(M\)

\[ \operatorname{ADD}=\frac1{|M|}\sum_{x\in M} \|\widehat R x+\widehat t-(Rx+t)\|_2, \]
\[ \operatorname{ADD\mbox{-}S}=\frac1{|M|}\sum_{x\in M} \min_{y\in M}\|\widehat R x+\widehat t-(Ry+t)\|_2. \]

ADD-S 只用于预声明的对称物体;最近邻会让非对称错误看起来过小。现代 symmetry-aware pose comparison 优先参考 BOP 的 MSSD/MSPD/VSD 及对象 symmetry 规范,而不是事后任选 ADD/ADD-S。

若方法真的重建 geometry,设预测/GT surface samples 为 \(P,Q\)

\[ \operatorname{CD}_{sym}(P,Q)= \frac1{|P|}\sum_{p\in P}\min_{q\in Q}\|p-q\|_2^2+ \frac1{|Q|}\sum_{q\in Q}\min_{p\in P}\|q-p\|_2^2, \]
\[ F_\tau=\frac{2P_\tau R_\tau}{P_\tau+R_\tau}, \]

其中 \(P_\tau,R_\tau\) 是两方向 surface distance \(\leq\tau\) 的比例。 ArtHOI hand--articulated reconstruction 使用 CD/MSSD(mm) 与 F-score@5/10mm,但 task 是 hand + articulated object,且某些 baseline 需要 pre-scan;只能在重叠输出、 同输入条件下比较。若 asset mesh 是所有方法共同输入,geometry 是常量,不应列作性能指标。

Metric 优点 风险 位置
(e_t,e_R) 直接、可解释 symmetry 与坐标约定 D3D main
ADD / MSSD 联合 pose+shape 表面影响 correspondence/symmetry 规则 GT mesh subtask main/secondary
ADD-S symmetry-aware 候选 nearest-neighbor 可低估错误 appendix,只对声明 symmetry
CD / F-score geometry fidelity ICP alignment/采样/阈值可投机 geometry-predicting methods only
scale/dimension 量纲与尺寸 scalar/anisotropic 定义混淆 D3D main,按原协议 cm

6. Articulation metrics

joint type、axis/origin 仅在方法实际预测时评测;若从给定 URDF/CAD 读取,不得计作性能。

\[ \operatorname{MAE}_{rev}=\frac1{|\mathcal I_r|}\sum_{(t,k)\in\mathcal I_r} d_r(\widehat q_{tk},q_{tk})\;[\mathrm{rad\ or\ deg}], \]
\[ \operatorname{MAE}_{pri}=\frac1{|\mathcal I_p|}\sum_{(t,k)\in\mathcal I_p} |\widehat q_{tk}-q_{tk}|\;[\mathrm m\ or\ cm], \]

RMSE 同样分表。对于单位 axis (a,\widehat a) 和轴线上 origin (o,\widehat o):

\[ e_{axis}=\cos^{-1}(|a^\top\widehat a|),\qquad e_{origin}=\|(\widehat o-o)\times a\|_2. \]

active-part endpoint / part-relative pose 更接近视觉后果:

\[ e_{link}=\frac1{|V_a|}\sum_{x\in V_a} \|\widehat{FK}(\widehat q)x-FK(q)x\|_2. \]
Metric 数据要求 优点 风险 位置
joint type macro-F1 type GT、方法预测 type 检验结构发现 class imbalance subtask appendix
axis/origin calibrated joint GT 诊断 kinematic discovery axis sign/origin gauge subtask appendix
q MAE/RMSE q GT 最直接 rad/m 混合、angle wrap D3D core;ARCTIC 仅激活时,分 type
endpoint/part pose mesh+FK GT 体现真实表面误差 geometry/scale 混入 key secondary
temporal q error aligned q trajectory/FPS 揭示抖动/相位 alignment 可投机 secondary

ARCTIC 的 AAE 为 \(T^{-1}\sum_t|\omega_t-\widehat\omega_t|\),表中单位 degree;D3D 对 revolute/prismatic 分别用 degree/cm。代码中的项目自定义 open-state threshold max(5 degree or 2 cm, 0.2 * max GT excursion) 只能作为 diagnostic/sensitivity,不能冒充标准。

7. Contact、penetration、support

有独立 contact GT

对 predicted/GT contact sets 计算:

\[ P=\frac{TP}{TP+FP},\quad R=\frac{TP}{TP+FN},\quad F_1=\frac{2PR}{P+R}. \]

必须固定 surface sampling、距离阈值、hand/part identity 与 frame tolerance,并报告 label coverage 和 wrong-part/wrong-hand contact。但当前 Recon output contract 没有 predicted contact set,因此 P/R/F1 当前是 Unavailable;只有另建 contact-prediction head/task、使用独立 GT 且 test 无人工 intervention 时才能作 secondary。

ARCTIC 对每帧、每只手 \(h\in\{l,r\}\) 在 GT 中距离 \(<3\) mm 的 hand--object vertex pairs \(\mathcal C_{t,h}^{GT}\) 定义:

\[ \operatorname{CDev}_{t,h}=\frac1{|\mathcal C_{t,h}^{GT}|} \sum_{(i,j)\in\mathcal C_{t,h}^{GT}} \|\widehat h_{ti}-\widehat o_{tj}\|_2\;[\mathrm{mm}], \qquad \operatorname{CDev}_{t,ho}= \operatorname{nanmean}_{h\in\{l,r\}}\operatorname{CDev}_{t,h}. \]

即先分手在 contact vertices 内平均,再对该帧的可用左/右手做 equal-hand nanmean,不把两手所有 contact pairs 混成一个按 vertex count 加权的集合。之后按官方 valid-frame 口径聚合。对 GT 中距离全窗都 \(<3\) mm、且长度 \(n-m+1\ge15\) frames (30 FPS) 的最大 stable-contact window \(w=(i,j,m,n)\)

\[ \mathcal V_w=\{t\in[m+1,n]:t-1\text{ 与 }t\text{ 均 valid}\}, \qquad \operatorname{MDev}_w=\frac1{|\mathcal V_w|}\sum_{t\in\mathcal V_w} \left\|(\widehat h_{ti}-\widehat h_{t-1,i})- (\widehat o_{tj}-\widehat o_{t-1,j})\right\|_2\;[\mathrm{mm}], \]

\(|\mathcal V_w|=0\) 则该 window unavailable;全窗都 valid 时分母才简化为 \(n-m\)。 然后对两手的全部 valid windows 平均。它检查 GT-contact pair 在预测中是否同向运动,但不是 contact classification。官方 relative-root 指标为

\[ \operatorname{MRRPE}_{a\to b,t}= \left\|(J^{a}_{0,t}-J^{b}_{0,t})- (\widehat J^{a}_{0,t}-\widehat J^{b}_{0,t})\right\|_2\;[\mathrm{mm}], \]

并分开报 \(r\to l\)\(r\to o\),不事后平均成一列。对 entity \(e\in\{h,o\}\) 的 root-relative vertices \(x^e_{tv}\),官方 30 FPS acceleration 为

\[ u^e_{tv}=\frac{x^e_{t-1,v}-2x^e_{t,v}+x^e_{t+1,v}}{(1/30)^2}, \qquad \operatorname{ACC}_{e}=\frac1{|\mathcal T^v_e||V_e|} \sum_{t\in\mathcal T^v_e,v} \|\widehat u^e_{tv}-u^e_{tv}\|_2\;[\mathrm{m/s^2}], \]

其中 \(\mathrm{ACC}_h\) 按官方规则平均左/右手,\(\mathrm{ACC}_o\) 单列,只使用三帧均 valid 的 centered stencil。ARCTIC object vertex success 先各自减掉 object base-center root:

\[ \operatorname{SR}_{0.05,t}=\frac{100}{V_o} \sum_{v=1}^{V_o}\mathbf1\!\left[ \left\|(\widehat o_{tv}-\widehat r^o_t)-(o_{tv}-r^o_t)\right\|_2 <0.05D_o\right]. \]

这是顶点通过比例 [%],不是整个 sequence 的 binary success;\(D_o\) 是 object diameter。ARCTIC official paperofficial code固定了上述定义; ContactOpt 展示 contact precision/recall 需要独立 thermal contact GT,而非仅靠最近距离。

没有 contact GT

D3D 只能报告 label-qualified gap distribution、wrong-part gap、active-label coverage 和可视化审计;mask 或 contact label 若参与优化,同一信号上的 IoU/gap 是 input consistency,不是独立 accuracy。

penetration depth/volume 需要经过 watertight、SDF sign 和 resolution audit 的 mesh/collision proxy;否则为 N/A。support/float 需冻结 ground/support geometry、gravity 与接触力定义。只在 penetration frames 或成功 sequences 上平均会产生条件分母偏差,必须同时报告全 ITT incidence 与 conditional magnitude。

Metric 位置
contact P/R/F1、wrong-part 当前 unavailable;独立 prediction+GT task 成立后 secondary
CDev/MDev/MRRPE/ACC/object SR ARCTIC official panel(仅激活时);需冻结 MANO↔SMPL-X bridge
label-qualified gap D3D diagnostic
penetration incidence + depth watertight audit 后 appendix/secondary
support/float downstream Physics diagnostic

8. Temporal metrics

对位置/vertex \(x_t\)\(\Delta t=1/\mathrm{FPS}\)

\[ v_t=\frac{x_{t+1}-x_t}{\Delta t},\quad a_t=\frac{x_{t+1}-2x_t+x_{t-1}}{\Delta t^2},\quad j_t=\frac{a_{t+1}-a_t}{\Delta t}. \]

旋转用 \(SO(3)\) log/geodesic increments,不能直接差分 Euler/axis-angle 参数。区分:

  • 与 GT 的 velocity/acceleration/jerk error;
  • prediction-only smoothness(只支持 regularity,不支持 accuracy);
  • long-horizon human-root/object drift;
  • phase timing/consistency;
  • foot slide(只在 GT/可靠 foot contact window);
  • valid duration/coverage。

对冻结的 foot-contact mask \(c_{t,f}\)、地面水平投影 \(\Pi_g\) 与 foot point \(x_{t,f}\),候选 foot slide 是每个连续 contact transition 的平均水平位移:

\[ \operatorname{FS}=\frac{1000}{|\mathcal K|} \sum_{(t,f)\in\mathcal K} \|\Pi_g(x_{t+1,f}-x_{t,f})\|_2\;[\mathrm{mm}], \quad \mathcal K=\{(t,f):c_{t,f}=c_{t+1,f}=1\}. \]

必须披露 foot points、contact detector/GT provenance、ground plane、FPS 和 \(|\mathcal K|\);空集合 为 Unavailable。这是 WHAM 风格的 prediction-only control-quality appendix 量,单位是 mm/步平均位移,不写成 mm/s;它不能代替 pose accuracy。

ARCTIC ACC 对 root-relative vertices 用 30 FPS centered difference 并除以 Δt²,单位 m/s²。当前项目 直接 np.diff(..., n=2/3) 的 raw differences 没除 FPS,只能叫 raw temporal difference diagnostic, 不得标 acceleration/jerk SI unit。

9. Generative / perceptual metrics

HumanML3D / text-to-motion 的指标只适用于冻结 learned feature extractor \(\phi_m,\phi_t\)、相同 prompt/reference distribution、完整 scheduled sample 及(Multimodality 时)每 prompt 多个 stochastic samples。

\[ \operatorname{FID}=\|\mu_r-\mu_g\|_2^2+ \operatorname{Tr}\!\left(\Sigma_r+\Sigma_g- 2(\Sigma_r\Sigma_g)^{1/2}\right)\downarrow, \]
\[ \operatorname{RPrec@k}=\frac1N\sum_i \mathbf1[i\in\operatorname{TopK}_{j}\{-\|\phi_m(x_i)-\phi_t(c_j)\|_2\}]\uparrow, \]
\[ \operatorname{Diversity}=\frac1L\sum_{\ell=1}^{L} \|\phi_m(x_{a_\ell})-\phi_m(x_{b_\ell})\|_2\uparrow, \quad a_\ell\ne b_\ell, \]
\[ \operatorname{Multimodality}=\frac1{PL} \sum_{p=1}^{P}\sum_{\ell=1}^{L} \|\phi_m(x_{p,a_\ell})-\phi_m(x_{p,b_\ell})\|_2\uparrow. \]

FID 与 feature distances 是 embedding-dependent 无量纲 score,R-Precision 是比例/[%];需报 extractor revision、reference/generated N、prompt 数、samples/prompt、pair draw seed 和 repeats。 user study 报预注册问题的 preference rate \(\widehat p=N^{-1}\sum_n\mathbf1[y_n=\mathrm{ours}]\) 与按 participant 及 item 聚类的 CI,并披露 ties、randomization、blinding、exclusion、IRB/ethics 与 sample-size rationale。

Metric 数据要求 / 优点 可投机/限制 结论
FID 大样本 reference/generated embeddings;分布级对照 选特征、小 N、缩放可改 ranking 有 generation claim 才 appendix;不进 Recon main
R-Precision paired prompt+sample;语义 retrieval negative-pool 和 encoder 敏感 同上;不证明 object/q/contact
Diversity 跨 sample 两两 feature distance 高噪声也可提高 同上;不支持单 case fidelity
Multimodality 每 prompt 多样本 deterministic Recon 无法定义 只进 stochastic generation appendix
User study 盲法人类评审;可评自然/语义 顺序、参与者与样例选择偏差 只作 perceptual appendix;不替代 GT/Physics

因此它们只在我们明确提出生成/感知质量 claim 且 protocol 合法时进入 appendix;不进 Recon accuracy 或 Physics 主表。

10. Full-pipeline downstream / controlled evidence

正文 D3D end-to-end comparison 固定同一个 frozen Ours Recon reference,让 CoDA、InterMimic、RePHO 与 Ours Physics backend 接收相同 case、asset、budget 和单次 rollout 合同。该表证明真实 reconstructed input 下的 backend difference;不把只有 RePHO 才拥有、且仍为 rigid-object task 的 native reconstruction pipeline 强行并入 articulated complete-pipeline ranking。

Ours Recon (kinematic / w/o Physics) 与 Full Ours 的 paired stage ablation 使用同一 reference,但只比较 两行都定义良好的 reconstruction、interaction 与 physics-aware metrics;kinematic row 不填被动物理 execution 的 outcome success / full-horizon success。ARCTIC/ParaHome GT references 另表绕过 Recon,作为 controller control。 Controlled corruption 只有在最终明确提出某类 noise robustness claim 时才追加。

11. Baseline 与状态

方法来源、最小 baseline 集合和 target lifecycle 只在 Reconstruction baselines 维护。本页不复制第二张状态表。

结论是:D3D-HOI official 与 Ours 是 Table 1 的必做 D3D rows;Ours w/o HOI Align. 是 Table 4/A3 的 paired ablation。3DADN 只有在 逐视频预测统一重评成功后才作为 Rev-only row; ARCTIC 空表保留 ArcticNet-SF + HOPformer + Ours;optional HOI-alignment ablation 独立成 panel,LSTM 只服务 temporal claim。 完整 3DHOI 九指标协议、JointTransformer、ArtHOI/Hand-ArtHOI 与 GRAIL 留在 related work、appendix、diagnostic 或 reported-only。

12. Provisional metric recommendation

Panel Main / lead Key secondary Diagnostic / appendix
D3D Real RGB · P0 core native object orientation[deg]、location[cm]、dimension[cm]、part motion(rev deg / pri cm 分列) endpoint active-link error(有 GT 时);ITT failure/coverage 单列 audit input-mask consistency、label-qualified gap、raw temporal difference
ARCTIC RGB official hands+object · P1 time-gated \(\mathrm{MPJPE}_h\)[mm]、\(\mathrm{MRRPE}_{r\to l}/\mathrm{MRRPE}_{r\to o}\)[mm]、AAE[deg]、object vertex [email protected][%]、CDev/MDev[mm]、\(\mathrm{ACC}_h/\mathrm{ACC}_o\)[m/s²] root metrics;run failure/coverage 单列 audit;contact P/R/F1 当前 unavailable PA-MPJPE、per-object、hand groups;activation gate 不过整块删除
ARCTIC RGB SMPL-X full-body supplement optional compact: world/camera MPJPE[mm]、root translation[mm]、coverage/failure root-aligned MPJPE/PVE、root orientation、velocity/ACC、foot slide 只有 camera/world、SMPL-X regressor、visibility 与 current configured SAM-Body4D dependency/lifecycle gate 全过才入表;否则删除并收缩 body-accuracy claim
Generated Video 有 frozen video reference 时同对应 accuracy;否则 failure/coverage + downstream Physics reprojection/input consistency(明确自洽) generative/perceptual only if claim
Full pipeline 唯一 confirmatory primary:ITT full-horizon success shared-Ours-Recon backend comparison、w/o Physics、q/contact error--success correlation;claim-triggered corruption curves

阈值(contact distance、F-score、object vertex success、geometry tolerance)保留原论文值作 reported comparison;我们的 paper threshold 必须按 annotation repeatability、measurement noise、visible validation 与 blinded human agreement 冻结,不能选择使 ours 获益最大的数值。

13. 待人确认与复审 gate

不阻塞文档完成、但在 formal freeze 前必须由人签名:

  1. confirmatory Physics primary 的最终 outcome/wrapper/threshold;
  2. D3D exact denominator(256/270/eligible)和 formal split;
  3. ARCTIC 是否整体激活:view/protocol、formal official test/eval server 或 untouched subject holdout、 MANO↔SMPL-X bridge、ArcticNet-SF、HOPformer、Ours overlap 任一不可得则整块不进 formal queue;
  4. generated video 是否实际导出逐帧可控 reference;
  5. 若 ARCTIC 激活,是否保留 temporal claim;保留才加 ArcticNet-LSTM。JointTransformer 只在 exact challenge code 同协议复现成功时进 appendix;
  6. contact/F-score/geometry 阈值与 practically meaningful effect margin;
  7. 能否取得或重新生成 3DADN 的 239-video 逐视频预测,并按 D3D native q-state 与统一帧集合重评;不能则该行降为 appendix reported-only;
  8. ArtHOI 是否存在严格匹配的 Generated protocol;完整 3DHOI 九指标协议不进入本项目 D3D main。

进入 private formal test 的 gate:输出 contract 与代码一致;panel availability matrix 签名;formal baseline 精确命名;manifest-driven ITT 和真实 CAD cluster 聚合跑通;所有公开数字 target 可复核;P0=0。