KIT | KIT-Bibliothek | Impressum | Datenschutz

How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models

Corbo, Simone; Bancale, Luca; Gennaro, Valeria De; Lestingi, Livia; Scotti, Vincenzo ORCID iD icon 1; Camilli, Matteo
1 Institut für Informationssicherheit und Verlässlichkeit (KASTEL), Karlsruher Institut für Technologie (KIT)

Abstract:

This contribution is an extended abstract of the paper originally published in the IEEE Transactions on Software Engineering (TSE) [Co25]. The paper presents EvoTox, an automated black-box testing framework that uses evolutionary search to assess Large Language Models’ susceptibility to generating toxic content through natural, realistic prompts. Empirical evaluation on five state-of-the-art LLMs (7-671B parameters) shows that EvoTox significantly outperforms existing baseline methods in detecting toxicity with effect sizes up to 1.0, while maintaining limited cost overhead (22−35%) and generating human-like prompts validated by domain experts


Zugehörige Institution(en) am KIT Institut für Informationssicherheit und Verlässlichkeit (KASTEL)
Publikationstyp Proceedingsbeitrag
Publikationsjahr 2026
Sprache Englisch
Identifikator ISSN: 2944-7682, 1617-5468
KITopen-ID: 1000196047
Erschienen in Software Engineering 2026 - Proceedings
Veranstaltung Software Engineering (SE 2026), Bern, Schweiz, 23.02.2026 – 27.02.2026
Verlag Gesellschaft für Informatik (GI)
Seiten 31–32
Serie Lecture Notes in Informatics (LNI), Proceedings - Series of the Gesellschaft fur Informatik (GI) ; P-377
Schlagwörter Automated Testing, Evolutionary Testing, Large Language Models, Toxic Speech
Nachgewiesen in Scopus
Relationen in KITopen
KIT – Die Universität in der Helmholtz-Gemeinschaft
KITopen Landing Page