Benchmarking Hallucination Detection in LLMs for Regulatory Applications Using SelfCheckGPT
Keywords:
hallucination detection, large language models, SelfCheckGPT, TruthfulQA, regulatory compliance, factual accuracyAbstract
In highly regulated sectors, SelfCheckGPT and TruthfulQA evaluate hallucination detection techniques in large language models (LLMs). We evaluate GPT-4, Claude, and LLaMA-3 to see whether automated hallucination identification may minimize factual mistakes and regulatory misinterpretations in compliance-intensive businesses including clinical trials, banking, and insurance Assessing compliance with industry-specific checklists and regulatory criteria may quantify model output deviations. Domain-specific language, nested phrases, and context-dependent facts impact intrinsic and extrinsic hallucination detection. SelfCheckGPT's multi-step approach susceptibility verification beats TruthfulQA's adversarial questioning. Enterprise-level AI governance systems should identify hallucinations to reduce factual and legal mistakes. This work examines domain-aware regulatory LLM reliability.
Downloads
References
M. Manakul, Y. L. Lertvittayakumjorn, and A. Vlachos, “SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models,” arXiv preprint arXiv:2303.08896, 2023.
S. Lin, M. Hilton, H. Evans, and D. Song, “TruthfulQA: Measuring how models mimic human falsehoods,” Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022, pp. 3214–3234.
T. Jiang, J. R. Klinger, and P. Liang, “Assessing the factual accuracy of large language models through model self-evaluation,” arXiv preprint arXiv:2305.13243, 2023.
J. Maynez, S. Narayan, B. Bohnet, and R. McDonald, “On faithfulness and factuality in abstractive summarization,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020, pp. 1906–1919.
S. Ji, S. Lee, and H. Kang, “Survey of hallucination in natural language generation: Types, detection, and mitigation,” ACM Computing Surveys, vol. 56, no. 3, pp. 1–38, 2024.
T. Kotonya and F. Toni, “Explainable automated fact-checking: A survey,” Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), 2021, pp. 4722–4729.
Y. Zhang, D. Liu, and M. Zhou, “Benchmarking truthfulness and factual consistency in large language models,” Transactions on Machine Learning Research (TMLR), 2024.
OpenAI, “GPT-4 Technical Report,” arXiv preprint arXiv:2303.08774, 2023.
Anthropic, “Claude 3 model card: Safe and reliable reasoning in aligned large language models,” Anthropic Research Blog, 2024.
Meta AI, “LLaMA 3: Open foundation and fine-tuned chat models,” Meta AI Research Report, 2024.
A. Tamkin, J. Perez, and D. Ganguli, “Understanding the capabilities and limitations of LLM self-verification,” Neural Information Processing Systems (NeurIPS) Workshop on Trustworthy and Socially Responsible Language Models, 2023.
E. Gabrilovich, A. Fabbri, and K. Duh, “Evaluating factual consistency in generative AI systems: Metrics, benchmarks, and best practices,” AI Magazine, vol. 45, no. 2, pp. 87–104, 2024.
C. Chen, L. Chen, and Y. Wang, “Automated detection of hallucinations in legal LLM outputs using retrieval-augmented reasoning,” Proceedings of the 2024 International Conference on Computational Linguistics (COLING), pp. 1189–1203.
D. Xu, J. Wang, and F. Wu, “Fact-aware decoding for mitigating hallucination in large language models,” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1347–1361.
European Union, “EU Artificial Intelligence Act: Proposal for a regulation of the European Parliament and of the Council laying down harmonised rules on artificial intelligence,” Official Journal of the European Union, COM(2021) 206 final, 2021.
S. Bommasani et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021.
P. Mitchell, L. Raifer, and R. Danks, “Factual reliability and epistemic calibration in generative AI for compliance systems,” Journal of Artificial Intelligence Research, vol. 77, pp. 112–138, 2024.
M. Shuster, A. Roller, and J. Weston, “The limitations of large language models in knowledge-intensive tasks,” Proceedings of the 37th AAAI Conference on Artificial Intelligence, 2023, pp. 14211–14221.
M. Ribeiro, S. Singh, and C. Guestrin, “Anchors: High-precision model-agnostic explanations,” Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI), 2018, pp. 1527–1535.
R. Leike, D. Krueger, and P. Christiano, “Scalable oversight for large language models via self-consistency and debate,” arXiv preprint arXiv:2306.03341, 2023.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.