Computer Science > Artificial Intelligence

arXiv:2311.12320 (cs)

[Submitted on 21 Nov 2023]

Title:A Survey on Multimodal Large Language Models for Autonomous Driving

View PDF

Abstract:With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and control tools as humans. In recent months, LLMs have shown widespread attention in autonomous driving and map systems. Despite its immense potential, there is still a lack of a comprehensive understanding of key challenges, opportunities, and future endeavors to apply in LLM driving systems. In this paper, we present a systematic investigation in this field. We first introduce the background of Multimodal Large Language Models (MLLMs), the multimodal models development using LLMs, and the history of autonomous driving. Then, we overview existing MLLM tools for driving, transportation, and map systems together with existing datasets and benchmarks. Moreover, we summarized the works in The 1st WACV Workshop on Large Language and Vision Models for Autonomous Driving (LLVM-AD), which is the first workshop of its kind regarding LLMs in autonomous driving. To further promote the development of this field, we also discuss several important problems regarding using MLLMs in autonomous driving systems that need to be solved by both academia and industry.

Subjects:	Artificial Intelligence (cs.AI)
Cite as:	arXiv:2311.12320 [cs.AI]
	(or arXiv:2311.12320v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2311.12320

Submission history

From: Can Cui [view email]
[v1] Tue, 21 Nov 2023 03:32:01 UTC (18,131 KB)

Computer Science > Artificial Intelligence

Title:A Survey on Multimodal Large Language Models for Autonomous Driving

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:A Survey on Multimodal Large Language Models for Autonomous Driving

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators