vmr3D / docs /ROADMAP.md
redazul's picture
Add VMR research artifacts, validated animation and learning roadmap
1ef5660 verified
|
Raw History Blame Contribute Delete
2.28 kB

VMR learning roadmap

Status: proposed work. The existing evidence is a reconstruction case study and playable animation, not a trained general model.

  1. Make the teacher repeatable. Package reconstruction, fitting, corrections and export so another sequence can run with recorded inputs, parameters and human intervention. Report runtime and failure cases. Preserve the current accepted example as a regression case.
  2. Build varied supervised examples. Render known animated characters with exact geometry and cameras; add reviewed video reconstructions. Record camera conventions, stable correspondence, visibility, contact and correction provenance. Preserve rejected outputs with failure labels. Keep all derivatives of one source together in one evaluation split.
  3. Learn one bounded task. Given a reference character and a short video segment, predict motion controls and local corrections. Compare the learned initialization plus refinement against fitting alone, including runtime and accuracy. Start with a predictive model; generative uncertainty is optional at this stage.
  4. Evaluate generalization. Hold out entire character identities and source motions, not neighboring frames. Measure surface fit, distortion, temporal consistency and contact preservation separately. Use blinded visual comparisons where feasible. Temporal scores must distinguish rapid intentional motion from jitter.
  5. Expand generation and export. Evaluate pretrained character generation, then learned latent motion generation where useful. Solve or explicitly represent topology changes. Reduce dense motion to editable animation curves within a documented error tolerance.
  6. Release a real checkpoint. Publish learned weights, configuration, preprocessing, runnable inference, training details and held-out evaluation when available. An adapted existing architecture is acceptable; a new architecture is not required.

A candidate architecture separates video identity/camera estimation, a reference mesh with deformation controls, temporal motion prediction and local correction, followed by validated export. The exact network, losses, data scale and training budget remain decisions to be justified by experiments. No generalization results are claimed now.