We use cookies to improve your experience with our site.
Xing Y, Huang JQ, Zhao HJ et al. Safety evaluation for large language models with metamorphic testing: An empirical study. JOURNAL OFCOMPUTER SCIENCE AND TECHNOLOGY, 41(3): 924−935, May 2026. DOI: 10.1007/s11390-026-5705-z
Citation: Xing Y, Huang JQ, Zhao HJ et al. Safety evaluation for large language models with metamorphic testing: An empirical study. JOURNAL OFCOMPUTER SCIENCE AND TECHNOLOGY, 41(3): 924−935, May 2026. DOI: 10.1007/s11390-026-5705-z

Safety Evaluation for Large Language Models with Metamorphic Testing: An Empirical Study

  • Although large language models (LLMs) have been widely deployed across numerous applications, they can generate harmful or illicit content, posing substantial safety risks. Evaluating such risks requires effective evaluation methodologies using high-quality benchmarking datasets. This study introduces LLMSafetyChoice in this regard, a multilingual benchmark for content safety evaluation containing 11911 multiple-choice questions in both Chinese and English, covering four safety domains and eight categories per language (nine in total). We further introduce a systematic metamorphic-testing approach, defining seven metamorphic relations for LLMSafetyChoice, for LLM safety evaluations. Through an extensive empirical study involving 1408 evaluation scenarios (11 LLMs×(8 categories×2 languages)×(1 constructed benchmark + 7 transformations)), we reveal key insights into model behavior under safety-critical conditions and demonstrate that metamorphic testing effectively uncovers subtle safety vulnerabilities. The benchmark and evaluation results are publicly available at https://anonymous.4open.science/r/LLMMetamorphic-08C9/.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return