Towards an AI-Enabled Metaverse: Architecture, Resource Allocation, and Multimodal Benchmarking
| dc.contributor.author | Long, Zijian | |
| dc.contributor.supervisor | El Saddik, Abdulmotaleb | |
| dc.contributor.supervisor | Dong, Haiwei | |
| dc.date.accessioned | 2026-07-16T13:14:52Z | |
| dc.date.issued | 2026-07-16 | |
| dc.description.abstract | The metaverse, envisioned as a paradigm for next-generation digital environments, remains fundamentally constrained by isolated, pre-defined, and reactive virtual systems. Artificial intelligence (AI) provides a conceptual and technical foundation for overcoming these limitations by enabling persistent perception, autonomous decision-making, and adaptive coordination. This thesis investigates how AI can be systematically embedded into metaverse systems from a system-level perspective. It begins with a comprehensive analysis of metaverse network traffic characteristics, examining the necessity, performance implications, and trade-offs between remote rendering and local rendering. Building on these empirical insights, the thesis proposes a unified thing–edge–cloud architecture that coordinates heterogeneous intelligent agents operating across network layers. Within this architecture, the thesis addresses two core challenges in AI-enabled metaverse systems: adaptive resource allocation and cognitive scene understanding. At the edge layer, adaptive bandwidth allocation for immersive streaming is formulated as a cooperative multi-agent decision-making problem under dynamic network conditions. A Multi-Agent Soft Actor Critic (MASAC)–based strategy is developed and yields consistent improvements in user Quality of Experience (QoE) of at least 14\% compared with representative streaming baselines. At the cloud layer, the thesis introduces a metaverse-oriented benchmarking framework for evaluating the scene understanding capabilities of Multimodal Large Language Models (MLLMs). Thirteen representative MLLMs are assessed through a pairwise comparison protocol under an MLLM-as-a-judge paradigm guided by metaverse-specific semantic criteria. Human expert validation confirms the robustness of the proposed benchmark, achieving an agreement rate of 87.6\% with human consensus. Overall, this thesis establishes an integrated methodological and architectural foundation for the design, optimization, and evaluation of AI-enabled metaverse systems. | |
| dc.identifier.uri | http://hdl.handle.net/10393/51853 | |
| dc.identifier.uri | https://doi.org/10.20381/ruor-32092 | |
| dc.language.iso | en | |
| dc.publisher | Université d'Ottawa | University of Ottawa | |
| dc.subject | Metaverse | |
| dc.subject | Edge resource allocation | |
| dc.subject | Large language model | |
| dc.subject | Deep reinforcement learning | |
| dc.title | Towards an AI-Enabled Metaverse: Architecture, Resource Allocation, and Multimodal Benchmarking | |
| dc.type | Thesis | en |
| thesis.degree.discipline | Génie / Engineering | |
| thesis.degree.level | Doctoral | |
| thesis.degree.name | PhD | |
| uottawa.department | Science informatique et génie électrique / Electrical Engineering and Computer Science |
Fichiers
Trousse originale
1 - 1 sur 1
En cours de chargement...
- Nom:
- Long_Zijian_2026_thesis.pdf
- Taille:
- 5.78 MB
- Format:
- Adobe Portable Document Format
Trousse de licence
1 - 1 sur 1
En cours de chargement...
- Nom:
- license.txt
- Taille:
- 2.51 KB
- Format:
- Item-specific license agreed upon to submission
- Description:
