Yujin Chen 陈雨劲
I am a final-year Ph.D. candidate in the Visual Computing Lab at the Technical University of Munich, supervised by Prof. Matthias Nießner. I am currently a Research Scientist Intern at Meta Reality Labs in Zurich. I received my B.Eng. and M.Sc. from Wuhan University.
I am seeking a full-time position, with an anticipated start in early 2027.
My research focuses on learning to represent, understand, and generate the dynamic 3D world. My work spans 3D appearance modeling, 3D/4D representation learning, and human interaction. I am interested in developing, adapting, and applying foundation models for spatial intelligence and generative world modeling.
Experience
Research Scientist Intern, Meta Reality Labs, Zurich, Switzerland, Aug. 2026 – present
Research Scientist Intern, Meta Reality Labs, Redmond, United States, Jul. 2025 – Nov. 2025
Research Intern, Tencent AI Lab, Shenzhen, China, Dec. 2019 – Jun. 2021
Visiting Researcher, State University of New York at Buffalo, Buffalo, United States, Jul. 2019 – Nov. 2019
Research Assistant, Wuhan University, Wuhan, China, Jan. 2017 – Jun. 2021
Publications
Seen2Scene: Completing Realistic 3D Scenes with Visibility-Guided Flow
European Conference on Computer Vision (ECCV), 2026
Training sparse 3D transformers directly on incomplete real-world scans, with support for layout-, text-, and partial-scan conditioning.
EgoMAN: Interaction-Structured Reasoning for Egocentric 3D Hand Trajectory Prediction
European Conference on Computer Vision (ECCV), 2026
Aligning vision-language reasoning with motion generation through trajectory tokens and progressive training.
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors
Neural Information Processing Systems (NeurIPS), 2025
Adapting pretrained image priors through multi-view differentiable rendering and texture-space constraints, without additional model training.
Mesh2NeRF: Direct Mesh Supervision for Neural Radiance Field Representation and Generation
European Conference on Computer Vision (ECCV), 2024
Analytically deriving ground-truth radiance fields from textured meshes, then using them to supervise triplane-based diffusion models for conditional and unconditional 3D generation.
SSR-2D: Semantic 3D Scene Reconstruction from 2D Images
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2024
Using differentiable rendering to learn geometry completion, color, and semantics from 2D supervision, without 3D ground-truth annotations.
PHRIT: Parametric Hand Representation with Implicit Template
International Conference on Computer Vision (ICCV), 2023
Learning part-based signed distance fields and a skeleton-driven deformation field as a fully differentiable layer for downstream reconstruction tasks.
Consistent 3D Hand Reconstruction in Video via Self-supervised Learning
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2023
Enforcing temporal consistency in motion, shape, and texture during video training, using detected 2D keypoints and image appearance instead of 3D annotations.
4DContrast: Contrastive Learning with Dynamic Correspondences for 3D Scene Understanding
European Conference on Computer Vision (ECCV), 2022
Transferring motion-aware features learned from synthetic objects in real scans to 3D segmentation and detection, including settings with limited labeled data.
MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video
Conference on Computer Vision and Pattern Recognition (CVPR), 2022
Alternating spatial and temporal transformer attention to capture relationships between joints and long-range motion in a sequence-to-sequence architecture.
Model-based 3D Hand Reconstruction via Self-Supervised Learning
Conference on Computer Vision and Pattern Recognition (CVPR), 2021
Jointly estimating pose, shape, texture, and camera parameters using differentiable rendering and detected 2D keypoints, without 3D annotations.
I2UV-HandNet: Image-to-UV Prediction Network for Accurate and High-fidelity 3D Hand Mesh Modeling
International Conference on Computer Vision (ICCV), 2021
Formulating dense 3D surface regression as image-to-image translation, with learned UV-space refinement for high-resolution geometry.
SO-HandNet: Self-Organizing Network for 3D Hand Pose Estimation with Semi-supervised Learning
International Conference on Computer Vision (ICCV), 2019
Pretraining a point cloud autoencoder on unlabeled hand scans and sharing its encoder with a pose regressor to reduce reliance on 3D pose labels.
Services
Workshops
1st Workshop on Generating Digital Twins from Images and Videos, ICCV 2025 Workshop, Co-organizer
Reviewing
CVPR, ECCV, ICCV, NeurIPS, AAAI, TPAMI, IJCV, TIP
Teaching
- Introduction to Deep Learning (IN2346), Head Teaching Assistant, 2022–2026. Course with more than 1,000 students.
- Machine Learning for 3D Geometry (IN2392), Teaching Assistant, Summer 2022
- Advanced Deep Learning for Computer Vision: Visual Computing (IN2390), Teaching Assistant, Winter 2021
Misc
Beyond academia, I enjoy exploring the world and capturing moments. Feel free to chat with me about sports 🧗 (I'm into bouldering and tennis), travel ✈️ , photography 📸 (my Flickr album), or any other interesting topics.











