Review: The Splog Detection Task and A Solution Based on Temporal and Link Properties

Authors: Yu-Ru Lin, Wen-Yen Chen, Xiaolin Shi, Richard Sia, Xiaodan Song, Yun Chi, Koji Hino, Hari Sundaram, Jun Tatemura and Belle Tseng
Year: 2006
Published in: The Fifteenth Text REtrieval Conference (TREC 2006) Proceedings
Importance: High


Spam blogs (splogs) have become a major problem in the increasingly popular blogosphere. Splogs are detrimental in that they corrupt the quality of information retrieved and they waste tremendous network and storage resources. We study several research issues in splog detection. First, in comparison to web spam and email spam, we identify some unique characteristics of splog. Second, we propose a new online task that captures the unique characteristics of splog, in addition to tasks based on the traditional IR evaluation framework. The new task introduces a novel time-sensitive detection evaluation to indicate how quickly a detector can identify splogs. Third, we propose a splog detection algorithm that combines traditional content features with temporal and link regularity features that are unique to blogs. Finally, we develop an annotation tool to generate ground truth on a sampled subset of the TREC-Blog dataset. We conducted experiments on both offline (traditional splog detection) and our proposed online splog detection task. Experiments based on the annotated ground truth set show excellent results on both offline and online splog detection tasks.

My Review

Authors suggest new method for detecting splogs based on online approach beside other offline approaches.

Methods for deceiving search engines by splogs:

  1. relevancy – via keyword stuffing
  2. popularity – via link farm
  3. recency – via frequent posts

Splog characteristic

  1. Machine-generated content
  2. No value addition
  3. Hidden agenda, usually economic goal

Number 2 and 1 are not good classified, their definition cover each other.

Authors claimed that Splogs are different from Web Spams since:

  1. Blogs’ dynamic content
  2. Non-endorsement links

Above two reasons are not sufficient for differing splogs from web spasm. Mistakenly or not, Authors mentioned to two reasons for differing blogs from other websites. Also above two characteristics are well known in some kind of websites such as news website that updated frequently and get user opinions for each news article.

As paper goes authors mention to these facts that splogs content sometimes are copied from non-spam blogs so current web spam detection can not detect them. More or less this claim is true but they should provide good reasons or statistical surveys for proving this claim.

Personally thought that splog can be detected into 4 ways:

  1. Copied contents. We can found them by employing text comparing algorithms.
  2. Repetitive content. By employing current web spam detection algorithms.
  3. Link ring. By employing link farms detection algorithms.
  4. Sping. By employing Spam ping detection algorithms
  5. Design of blog. Blog visitors many times found whether or not blog is spam by looking at design of blog. Currently there is no work on this issue. Any I like to work on it.

In the rest of paper authors demonstrate their detection method, they used 5 feature for detecting splogs which are:

  1. Tokenized URLs
  2. Blog and post titles
  3. Anchor text
  4. Blog homepage content
  5. Post content

There is nothing new with these feature that authors claimed before, and they do not clearly mention to their detection method difference.  or at least I did not figure it out.


4 Responses to Review: The Splog Detection Task and A Solution Based on Temporal and Link Properties

  1. […] Review: The Splog Detection Task and A Solution Based on Temporal and Link Properties […]

  2. vidy says:

    Hi Pedram

    I agree with your opinion after reading this paper.

    Frankly speaking, the way the paper is written is totally crap.

    I am not sure what the 10 authors contributed when they wrote this paper.

    It does not follow even simple standards that are used in a research paper 🙂

    Well the am not convinced with the idea of time based splog detection, although it may be useful, at times discarding the blog earlier in the lifecycle can be harmful, and what happens if there are false positives?

    The authors dont talk about it as well.

    However your idea seems good, can u email me in detail your idea about Design of Blog?


  3. vidy says:

    Hi Pedram

    I just checked which conference published this paper, and I came across this article which may be useful for our ALR project


Leave a Reply

Fill in your details below or click an icon to log in: Logo

You are commenting using your account. Log Out / Change )

Twitter picture

You are commenting using your Twitter account. Log Out / Change )

Facebook photo

You are commenting using your Facebook account. Log Out / Change )

Google+ photo

You are commenting using your Google+ account. Log Out / Change )

Connecting to %s

%d bloggers like this: