Autonomous driving paper index
Hallucinations in generative artificial intelligence and large language models: tests, datasets, detection and correction methods
One-line summary
More specifically, this review encompasses their definitions, underlying mechanisms, taxonomies, commonly used tests, and datasets for evaluating hallucinations.
Engineering notes
Key topics: autonomous driving, large language model. See the paper for implementation details and experimental results.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。
Original abstract
Abstract Generative Artificial Intelligence (GAI) and Large Language Models (LLMs) have demonstrated significant capabilities in generating human-like content; however, they exhibit a propensity to fabricate spurious information, a phenomenon often termed hallucination . This review paper provides a comprehensive overview of hallucinations in GAI and LLMs. More specifically, this review encompasses their definitions, underlying mechanisms, taxonomies, commonly used tests, and datasets for evaluating hallucinations. In addition, this review dives into intrinsic and extrinsic factors contributing to these inaccuracies, including limitations in model architectures, training data biases, and inference algorithms, as well as examines various detection strategies [e.g., post-hoc consistency checks, external fact-checking, contrastive learning, uncertainty calibration methods, and Retrieval-Augmented Generation (RAG)]. The review also synthesizes a range of correction and mitigation techniques, from proactive measures during training to hybrid approaches that combine detection and intervention. Finally, this review integrates qualitative assessments and comparative insights to delineate the impact of hallucinations on user trust and acceptability, and to shed light on current challenges and future research trends.
Links and sources
Need this topic turned into a technical roadmap?
Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.
Request B2B research
Comments