Author(s): Fishchuk, Vitalii (2023)
Abstract:
This thesis explores the effectiveness of adversarial attack methods in evading AI-text detection. Experimenting on three attack categories, prompt engineering, parameter tweaking, and character-level mutations, this research employs a mixed-methods approach to examine the effectiveness of such attacks with the recently released GPT-3.5 model. Results from this research reveal the low robustness of existing detectors towards practical and resource-efficient attack methods. The findings demonstrate how prompt engineering, parameter tweaking and character-level mutations can be exploited to evade detection effectively. Additionally, the study shows that detector algorithms struggle with the GPT-4 model and highlights the need for urgent improvement in existing detectors.
Document(s):
fishchuk_BA_eemcs.pdf