3D-MRL: Nested Multimodal 3D Representations via Matryoshka Representation Learning

1Universidade de São Paulo, 2Bowdoin College
The British Machine Vision Conference (BMVC) 2026

Abstract

Vision-Language Models align point clouds with image and text embeddings, enabling zero-shot recognition, retrieval, and open-vocabulary understanding of 3D shapes. Existing multimodal 3D pre-training methods produce fixed-dimensional embeddings, requiring separate models for different computational budgets.

We propose 3D Matryoshka Representation Learning (3D-MRL), a multimodal 3D pre-training framework based on Matryoshka Representation Learning. 3D-MRL learns nested 3D representations by aligning point clouds with frozen CLIP image and text embeddings while applying contrastive supervision across multiple embedding dimensions. The Matryoshka objective is applied only to the 3D encoder, allowing a single model to produce representations at different dimensionalities without retraining.

Experiments on the Objaverse-LVIS, ModelNet40, and ScanNet datasets show that 3D-MRL achieves competitive performance on zero-shot and few-shot 3D recognition tasks. In addition, the learned representations support retrieval across different embedding dimensions within a single model. On Objaverse-LVIS, 3D-MRL improves Top-1 accuracy from 46.8% to 50.9%. Retrieval experiments further show that different embedding dimensions yield varying levels of semantic and geometric specificity.

Method Overview

Overview of the 3D-MRL framework

Our 3D-MRL learns nested 3D representations by aligning point clouds with frozen CLIP image and text embeddings across multiple dimensionalities. Contrastive supervision is applied to each nested prefix, enabling a single 3D encoder to produce representations from compact to full dimensionality without retraining. At inference, embeddings can be truncated according to the available computational budget, supporting efficient zero-shot recognition and cross-modal retrieval.

Zero-shot 3D Shape Classification

Zero-shot 3D shape classification results.

3D-MRL enables zero-shot 3D recognition by directly comparing nested point-cloud embeddings with frozen CLIP category embeddings, without task-specific fine-tuning. Across Objaverse-LVIS, ModelNet40, and ScanObjectNN, the model achieves competitive performance, with particularly strong gains on the long-tailed Objaverse-LVIS benchmark, reaching 50.9% Top-1 accuracy while preserving the ability to operate at multiple embedding dimensions.

Point cloud-input 3D shape retrieval

Qualitative results of 3D-to-3D retrieval on Objaverse-LVIS dataset.

Text-input 3D shape retrieval

Qualitative results of text-to-3D retrieval on Objaverse-LVIS dataset.

More Qualitative Examples

3D-MRL with the Matryoshka cascade search strategy, exploiting the nested structure of the learned embeddings for efficient retrieval.

Qualitative results of 3D-MRL cascade retrieval.

Text-to-3D retrieval results on the Objaverse-LVIS dataset, showing the top-3 retrieved shapes for each text query.

Qualitative results of text-to-3D retrieval on Objaverse-LVIS dataset.

Image-to-3D retrieval results on the Objaverse-LVIS dataset, showing the top-2 retrieved shapes for each image query.

Qualitative results of Image-to-3D retrieval on Objaverse-LVIS dataset.

BibTeX


       @misc{lobo20263dmrlnestedmultimodal3d,
              title={3D-MRL: Nested Multimodal 3D Representations via Matryoshka Representation Learning}, 
              author={Márcus Lobo and Vitor Matias and Jeová Farias and Moacir Ponti},
              year={2026},
              eprint={2608.29285},
              archivePrefix={arXiv},
              primaryClass={cs.CV},
              url={https://arxiv.org/abs/2608.29285}, 
        }
      

Acknowledgments

This work was supported in part by the São Paulo Research Foundation (FAPESP), under grants #2024/09462-1 and #2026/01721-3, and by the Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq) under fellowship grant #315158/2023-9. The authors gratefully acknowledge the Center for Mathematical Sciences Applied to Industry (CeMEAI) for providing computational resources, funded by FAPESP under grant #2013/07375-0.