DATA-DRIVEN CONTENT-BASED WEB SPAM DETECTION: A SYSTEMATIC REVIEW OF DATA SCIENCE AND MACHINE LEARNING APPROACHES
Keywords:
Data science; content-based web spam detection; spamdexing; machine learning; feature engineering; web-page classification; systematic literature reviewAbstract
Web spam uses deceptive or misleading techniques that are misleading in order to change how a website ranks on the search engine result pages (SERPs) and make finding useful information more difficult. This paper provides a systematic literature review on the subject of web spam recognition via a data science perspective. The paper provides an insight into the manner in which the data of a website is represented through text, HTML, semantic information, website structure, and link-related signals. The discussion concerning the techniques of data processing, feature engineering, data classification, machine learning, ensemble learning, and deep learning is also carried on. The reviewed studies show that data-driven performance depends on the choice of datasets, feature spaces, learning algorithms, class balance, and evaluation measures such as accuracy, precision, recall, and F-measure. The article also summarizes the strengths and weaknesses of existing content-based and combined detection approaches and identifies continuing challenges, including evolving spam strategies, limited benchmark diversity, computational cost, model generalization, and adversarial behavior. By organizing the literature around data, features, models, and evaluation, the review positions content-spam detection as an applied data science problem and highlights directions for more adaptive and reliable web spam detection.












