Instructions to use AppliedIntuitionResearch/BuildRome with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Trellis
How to use AppliedIntuitionResearch/BuildRome with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Building Rome from a Single Image
Building Rome generates a complete metric 3D scene mesh — visible surfaces and plausible occluded geometry — from a single photograph, indoors and outdoors. It fine-tunes the TRELLIS.2 object generator to work on scenes: a monocular depth estimate (MoGe-3) lifts image features onto the observed surfaces, the scene is covered with chunks that grow with distance from the camera, and the chunks are generated one at a time with latent outpainting and stitched into one mesh.
⚠️ Research release. This model is released for research purposes under the terms below. It is not a production system and does not include any Applied Intuition proprietary data, product code, or production checkpoints.
Built with DINOv3.
Model Details
| Developed by | Applied Intuition — AI Research |
| Model type | LoRA adapters and depth-conditioning branches for the TRELLIS.2 sparse-structure (SS) and shape (SLat) flow transformers |
| Paper | Building Rome from a Single Image |
| Code | https://github.com/Applied-Intuition-Open-Source/BuildRome |
| Project page | https://build-rome.github.io/ |
| Base model | microsoft/TRELLIS.2-4B (MIT), downloaded separately |
| Training data | Indoor scenes from SAGE-10k and Infinigen, and about 4,000 synthetic outdoor scenes (see the paper) |
| License (weights) | CC BY-NC 4.0 |
| Contact | GitHub Issues on the code repo |
Checkpoints
The files hold only the fine-tuned parameters; the inference code adds them to the public TRELLIS.2 base weights.
| Checkpoint | Description | Size | Link |
|---|---|---|---|
ss_delta.safetensors |
Sparse-structure (occupancy) model: LoRA r32 adapters and depth-lift / clearance branches, 115M params. | 438 MiB | ss_delta.safetensors |
slat_delta.safetensors |
Shape model: LoRA r32 adapters and depth-lift branch, 66M params. | 251 MiB | slat_delta.safetensors |
SHA-256 checksums are in SHA256SUMS.
Intended Use & Limitations
Intended use: research on single-image 3D scene reconstruction and generation.
Out of scope / limitations:
- Outputs are geometry only (no texture). Metric scale and camera come from the MoGe-3 depth estimate, so depth errors carry into the scene.
- Occluded regions are plausible completions, not reconstructions of the true scene.
- Meshes are not guaranteed to be watertight and can contain holes, as with TRELLIS.2.
- Requires a CUDA GPU; a 98 m street scene (4 depth bands, 15 chunks) peaks at about 16 GiB and takes about 6 min on one A100.
Commercial use is not permitted under CC BY-NC 4.0.
How to Use
from build_rome import ScenePipeline
pipe = ScenePipeline.from_pretrained("AppliedIntuitionResearch/BuildRome")
scene = pipe("photo.jpg")
scene.save("results/photo") # scene.glb, layout_bands.json, layout_topview.png, spiral.mp4
Command line:
hf download AppliedIntuitionResearch/BuildRome --local-dir checkpoints
python inference/image_to_scene.py --image photo.jpg \
--ss-ckpt checkpoints/ss_delta.safetensors --slat-ckpt checkpoints/slat_delta.safetensors --out results/
The pipeline also downloads TRELLIS.2-4B, TRELLIS-image-large, DINOv3 ViT-L/16 (gated: accept its license on Hugging Face first) and MoGe-3. Full installation and usage instructions: see the GitHub repository.
License
The weights in this repository are released under CC BY-NC 4.0, non-commercial use only (full text in LICENSE). They are used together with separately downloaded models that keep their own terms:
- TRELLIS.2-4B and TRELLIS-image-large: MIT.
- DINOv3: the DINOv3 License. Built with DINOv3.
- MoGe-3: MIT.
Training data: SAGE-10k (Apache-2.0), Infinigen scenes (BSD-3-Clause generator), and synthetic outdoor scenes produced with the pipeline described in the paper. No training data is included in this repository.
Citation
@article{yenphraphai2026rome,
title = {Building Rome from a Single Image},
author = {Yenphraphai, Jiraphon and Li, Fang and Xu, Tianshuo and Meng, Depu and Herau, Quentin and Hu, Yihan and Yeh, Raymond A. and Zhan, Wei},
journal = {arXiv preprint arXiv:2610.08790},
year = {2026}
}
Acknowledgments
This model builds on TRELLIS.2 and TRELLIS (Microsoft, MIT), MoGe (Microsoft, MIT) and DINOv3 (Meta, DINOv3 License). Training uses scenes from SAGE-10k (NVIDIA) and Infinigen (Princeton). We thank the authors for making their work available.
Model tree for AppliedIntuitionResearch/BuildRome
Base model
microsoft/TRELLIS.2-4B