LaViDa-R1

Summary

LaViDa-R1 is the reasoning post-training continuation of LaViDa-O. It mixes SFT, online GRPO, and best-of- self-distillation while using answer forcing and partial-state tree search to obtain non-zero training signal on difficult prompts.

Official Artifacts

  • Preprint: arXiv 2602.14147
  • Official author publication entry: Shufan Li
  • Release caveat: no dedicated official project page, code repository, or checkpoint was verified at ingest time.

Role In The Wiki

LaViDa-R1 belongs at the intersection of diffusion-language inference dynamics and multimodal post-training. Its partial-state branching resembles candidate-search over a generative trajectory, but it does not model actions or environment transitions and should not be described as a world model.