Original Reddit post

Imagine recording someone plugging a cable into a socket. The hand tracker follows the reach and approach. It briefly loses the hand during insertion, then recovers after the cable is connected. Most of the recording looks excellent. My concern is that a short tracking gap can remove the part of a demonstration that matters most for learning the task. MEgoVista’s evaluation penalizes missed detections instead of simply excluding them from the pose score. That’s a useful starting point. But the same number of missed frames could have very different consequences depending on when they occur. The accompanying MEgo framework describes hand detection and temporal tracking before MANO-based 3D reconstruction. That makes coverage worth examining alongside the quality of the reconstructed motion. For a small team choosing a tracker, I’d add an evaluation breakdown by task phase: approach, contact and withdrawal. For each phase, report pose error, detection coverage and the longest tracking gap. That is an additional check I’d propose, rather than a plug-insertion result reported in the paper. The recording might still provide useful supervision for reaching toward the socket. I’d mark the insertion interval as uncertain and decide separately whether it was usable. For manipulation datasets, the timing of a failure deserves a place in the evaluation. https://arxiv.org/abs/2609.16684 submitted by /u/Klutzy_Cap8492

Originally posted by u/Klutzy_Cap8492 on r/ArtificialInteligence