AI-SPAMGUARD: A REPRODUCIBLE MULTI-TRANSFORMER AND EXPLAINABLE DECISION-FUSION FRAMEWORK FOR ROBUST DETECTION OF LLM-GENERATED OPINION SPAM

Authors

  • Muhammad Mustafa
  • Gohar Abbas
  • Asad Ullah
  • Aamir Aftab

Keywords:

AI-generated text detection, opinion spam, large language models, BERT, ELECTRA, XLNet, explainable AI, robustness, fake reviews, reproducibility.

Abstract

Large language models (LLMs) can generate persuasive online reviews at negligible marginal cost, raising a practical threat to review-platform integrity. This paper presents AI-SpamGuard, a reproducible multi-transformer evaluation and decision-fusion framework for detecting LLM-generated hotel opinion spam. The study reconstructs and extends the LLM-review benchmark of Liyanage et al. using validated human reviews together with GPT-3, LLaMA, and paraphrased AI-generated reviews. After data auditing, the usable corpus contains 800 authentic reviews, 800 GPT-3 reviews, 790 LLaMA reviews, and 800 paraphrased reviews. Four transformer detectors—BERT, SciBERT, ELECTRA, and XLNet—are fine-tuned for 10 epochs with three independent seeds. Native-tokenizer mean generated-class F1 reaches 99.02% on GPT-3 (BERT), 96.76% on LLaMA (XLNet), and 99.28% on paraphrased reviews (ELECTRA). A post-hoc majority-vote decision layer computed from synchronized saved predictions improves LLaMA F1 to 97.04% and paraphrased F1 to 99.30%, although it is slightly below BERT on GPT-3, demonstrating that fusion is generator-dependent rather than universally beneficial. Robustness tests reveal the central limitation: strong in-domain performance does not guarantee cross-generator transfer. For example, the strong extension detector obtains only 52.21% generated-F1 when trained on GPT-3 and tested on LLaMA, while unseen-hotel evaluation remains substantially stronger. Explainability is provided through the reproducible TF-IDF/Multinomial-Naive-Bayes baseline using additive lexical attribution and a locality-weighted LIME-style surrogate. The paper also reports aggregate confusion matrices, probability ROC-AUC for the saved baseline models, tokenizer ablation, leave-one-generator-out evaluation, and reproducibility constraints. The results support a practical conclusion: detector validation must emphasize generator shift and protocol fidelity, not only near-saturated in-domain accuracy.

Downloads

Published

2025-07-03

How to Cite

Muhammad Mustafa, Gohar Abbas, Asad Ullah, & Aamir Aftab. (2025). AI-SPAMGUARD: A REPRODUCIBLE MULTI-TRANSFORMER AND EXPLAINABLE DECISION-FUSION FRAMEWORK FOR ROBUST DETECTION OF LLM-GENERATED OPINION SPAM. Spectrum of Engineering Sciences, 3(7), 1781–1812. Retrieved from https://thesesjournal.com/index.php/1/article/view/3833