Skip to main navigation Skip to search Skip to main content

Generative models for multimodal 3D scene generation

  • Yao Wei*
  • *Corresponding author for this work

Research output: ThesisPhD Thesis - Research UT, graduation UT

394 Downloads (Pure)

Abstract

Three-dimensional (3D) scene models are essential for how humans and intelligent systems perceive, reason about, and interact with the world. From digital avatars and product design to indoor environments and urban landscapes, 3D data provides semantic and spatial context that 2D data alone cannot fully capture. However, acquiring detailed and realistic 3D scenes remains a major bottleneck in terms of accessibility and scalability. Conventional approaches are typically resource-expensive, labor-intensive and often not friendly for non-expert users. Powered by advances in both deep generative architectures and multimodal learning paradigms, generative models have emerged as a promising alternative. Therefore, the need for controllable, coherent, and functionally meaningful 3D scene generation from diverse input modalities defines the central problem of this dissertation: developing generative models that can synthesize realistic 3D scenes under multimodal supervision. We begin at the object level, tackling the challenging task of generating flexible and controllable 3D point clouds from a single image, especially for general and structurally complex objects such as buildings. Extending beyond individual objects, we investigate multi-object 3D scene generation with an emphasis on coherence and functional usability. Specifically, we explore how to incorporate global semantics, layout regularization, and human interaction priors into generative pipelines to ensure that the synthesized scenes are not only visually realistic but also functionally aligned with human-aware functionality.

Original languageEnglish
QualificationDoctor of Philosophy
Awarding Institution
  • University of Twente
  • Faculty of Geo-Information Science and Earth Observation
Supervisors/Advisors
  • Vosselman, George, Supervisor
  • Yang, Michael Ying, Supervisor
Award date3 Sept 2025
Publisher
Print ISBNs978-90-365-6769-5
Electronic ISBNs978-90-365-6770-1
DOIs
Publication statusPublished - 3 Sept 2025

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 11 - Sustainable Cities and Communities
    SDG 11 Sustainable Cities and Communities

Fingerprint

Dive into the research topics of 'Generative models for multimodal 3D scene generation'. Together they form a unique fingerprint.

Cite this