Part-aware 3D generation can produce parts that look complete but are not physically realizable: neighbouring parts interpenetrate, they have no valid contact surfaces to stay connected, and the assembly falls over in simulation. SNAP3D resolves the inter-part penetration, recovers a contact graph between neighbouring parts, and introduces parameterized connectors at their contact surfaces, then refines each connector’s placement, orientation and dimensions against feedback from physical simulation while preserving the generated geometry. Assemblies are released under gravity exactly as generated, and the exported parts can be printed on a desktop FDM printer and joined by hand without glue or fasteners.
@misc{tuan2026snap3d,title={SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image},author={Tuan, Yu-Rou and Tsui, Hao-Tang and Ugrinovic, Nicol\'as and Kitani, Kris and Ma, Xiaoxuan},year={2026},note={Preprint},}
ECCVW
RareOcc: Controllable 4D Occupancy and LiDAR Generation for Long-Tail Driving Scenes
Arthur Jakobsson, Yu-Rou Tuan, Leron Julian, and 8 more authors
In ECCV Workshop on Safe and Defensive Autonomous Driving, 2026
An end-to-end generative framework that synthesizes realistic 4D occupancy and LiDAR data for safety-critical, long-tail driving scenarios. We introduce an entity-centric BEV-to-occupancy method using grounded latents and semantic-Gaussian splatting to preserve rare-class agents during scene edits, and a unified editable bird’s-eye-view interface that integrates real crash records, LLM text specifications, and manual scene edits for controllable data synthesis.
@inproceedings{tuan2026rareocc,title={RareOcc: Controllable 4D Occupancy and LiDAR Generation for Long-Tail Driving Scenes},author={Jakobsson, Arthur and Tuan, Yu-Rou and Julian, Leron and Hunt, Shawn and Suzuki, Kenta and Tanaka, Shinya and Bandegi, Mahdi and Gloomis, Jay and Fujiyoshi, Hironobu and Kitani, Kris and Ichnowski, Jeffrey},booktitle={ECCV Workshop on Safe and Defensive Autonomous Driving},year={2026}}
CVPRW
StaDy4D: Towards Complete 4D Static-Dynamic Reconstruction with SIGMA
Hao-Tang Tsui*, Yu-Rou Tuan*, Ethan Lai, and 1 more author
In CVPR Workshop on Generative 3D Reconstruction, 2026
Selected for an Oral and received the Best Paper Award at the CVPR 2026 GenRecon3D Workshop (non-archival).
StaDy4D is a large-scale paired static-dynamic 4D benchmark with 9K sequences and 1.9M frames, enabling complete evaluation of dynamic-scene reconstruction in autonomous driving. We propose SIGMA, a test-time-adapted static-dynamic reconstruction method that reduces depth error by 41% over the strongest feed-forward baseline and achieves best-in-class point-cloud reconstruction, with zero-shot transfer to real-world benchmarks.
@inproceedings{tsui2026stady4d,title={StaDy4D: Towards Complete 4D Static-Dynamic Reconstruction with SIGMA},author={Tsui, Hao-Tang and Tuan, Yu-Rou and Lai, Ethan and Wang, Chen-Yu},booktitle={CVPR Workshop on Generative 3D Reconstruction},year={2026},}
Preprint
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
Jeremy Tien, Abishek Anand*, Yu-Rou Tuan*, and 3 more authors
An OS-level agent-safety benchmark with 82 computer-use tasks across three corrigibility scenarios, evaluating 12 frontier models on OSWorld-Verified. We identify widespread misalignment in benign settings, including 100% override rates and safety failures in spawned subagents, which accessed restricted credentials in 50% of trials.
@article{tien2026rogue,title={ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use},author={Tien, Jeremy and Anand, Abishek and Tuan, Yu-Rou and Shen, Yuchen and Kolter, J. Zico and Nayebi, Aran},journal={Under review},year={2026}}
2024
ECCV
TrajPrompt: Aligning Color Trajectory with Vision-Language Representations
Li-Wu Tsao, Hao-Tang Tsui, Yu-Rou Tuan, and 5 more authors
In European Conference on Computer Vision (ECCV), 2024
We use text and visual prompts to enhance pedestrian trajectory prediction, optimizing 2D position retrieval from embeddings for accurate and faster inference. Cross-modal learning improves ADE and FDE performance by over 35%.
@inproceedings{tsao2024trajprompt,title={TrajPrompt: Aligning Color Trajectory with Vision-Language Representations},author={Tsao, Li-Wu and Tsui, Hao-Tang and Tuan, Yu-Rou and Chen, Pei-Chi and Wang, Kuan-Lin and Wu, Jhih-Ciang and Shuai, Hong-Han and Cheng, Wen-Huang},booktitle={European Conference on Computer Vision (ECCV)},year={2024},}