Although large language models (LLMs) have been widely deployed across numerous applications, they can generate harmful or illicit content, posing substantial safety risks. Evaluating such risks requires effective evaluation methodologies using high-quality benchmarking datasets. This study introduces LLMSafetyChoice in this regard, a multilingual benchmark for content safety evaluation containing
11911 multiple-choice questions in both Chinese and English, covering four safety domains and eight categories per language (nine in total). We further introduce a systematic metamorphic-testing approach, defining seven metamorphic relations for LLMSafetyChoice, for LLM safety evaluations. Through an extensive empirical study involving
1408 evaluation scenarios (11 LLMs×(8 categories×2 languages)×(1 constructed benchmark + 7 transformations)), we reveal key insights into model behavior under safety-critical conditions and demonstrate that metamorphic testing effectively uncovers subtle safety vulnerabilities. The benchmark and evaluation results are publicly available at
https://anonymous.4open.science/r/LLMMetamorphic-08C9/.