|
ABSTRACT
Currently, most popular Web search engines adopt some link-based ranking methods such as PageRank. Driven by the huge potential benefit of improving rankings of Web pages, many tricks have been attempted to boost page rankings. The most common way, which is known as link spam, is to make up some artificially designed link structures. Detecting link spam effectively is a big challenge. In this article, we develop novel and effective detection methods for link spam target pages using page farms. The essential idea is intuitive: whether a page is the beneficiary of link spam is reflected by how it collects its PageRank score. Technically, how a target page collects its PageRank score is modeled by a page farm, which consists of pages contributing a major portion of the PageRank score of the target page. We propose two spamicity measures based on page farms. They can be used as an effective measure to check whether the pages are link spam target pages. An empirical study using a newly available real dataset strongly suggests that our method is effective. It outperforms the state-of-the-art methods like SpamRank and SpamMass in both precision and recall.
REFERENCES
Note: OCR errors may be found in this Reference List extracted from the full text article. ACM has opted to expose the complete List rather than only correct and linked references.
| |
1
|
Abou-Assaleh, T. and Das, T. 2007. Extension and propagation of manual and automatic Web spam scores. In Proceedings of the 3rd International Workshop on Adversarial Information Retrieval on the Web (AIRWeb'07). ACM.
|
| |
2
|
|
| |
3
|
Albert, R., Jeong, H., and Barabasi, A.-L. 1999. The diameter of the world wide Web. Nature 401, 130--131.
|
 |
4
|
|
| |
5
|
Becchetti, L., Castillo, C., Donato, D., Leonardi, S., and Baeza-Yates, R. 2006. Using rank propagation and probabilistic counting for link-based spam detection. In Proceedings of the Workshop on Web Mining and Web Usage Analysis (WebKDD06). ACM Press.
|
 |
6
|
András Benczúr , István Bíró , Károly Csalogány , Tamás Sarlós, Web spam detection via commercial intent analysis, Proceedings of the 3rd international workshop on Adversarial information retrieval on the web, May 08-08, 2007, Banff, Alberta, Canada
[doi> 10.1145/1244408.1244424]
|
| |
7
|
Benczur, A. A., Csalogany, K., Sarlos, T., and Uher, M. 2005. Spamrank: Fully automatic link spam detection. In Proceedings of the 1st International Workshop on Adversarial Information Retrieval on the Web (AIRWeb'05).
|
 |
8
|
|
 |
9
|
|
| |
10
|
Andrei Broder , Ravi Kumar , Farzin Maghoul , Prabhakar Raghavan , Sridhar Rajagopalan , Raymie Stata , Andrew Tomkins , Janet Wiener, Graph structure in the Web, Proceedings of the 9th international World Wide Web conference on Computer networks : the international journal of computer and telecommunications netowrking, p.309-320, June 2000, Amsterdam, The Netherlands
|
 |
11
|
Carlos Castillo , Debora Donato , Luca Becchetti , Paolo Boldi , Stefano Leonardi , Massimo Santini , Sebastiano Vigna, A reference collection for web spam, ACM SIGIR Forum, v.40 n.2, p.11-24, December 2006
[doi> 10.1145/1189702.1189703]
|
 |
12
|
Carlos Castillo , Debora Donato , Aristides Gionis , Vanessa Murdock , Fabrizio Silvestri, Know your neighbors: web spam detection using the web topology, Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval, July 23-27, 2007, Amsterdam, The Netherlands
[doi> 10.1145/1277741.1277814]
|
| |
13
|
Chien, S., Fetterly, D., Manasse, M., Najork, M., and Ntoulas, A. 2007. Microsoft silicon valley Web spam challenge entry. In Proceedings of the 3rd International Workshop on Adversarial Information Retrieval on the Web (AIRWeb'07). ACM.
|
| |
14
|
Cormack, G. V. 2007. Content-based Web spam detection. In Proceedings of the 3rd International Workshop on Adversarial Information Retrieval on the Web (AIRWeb'07). ACM.
|
| |
15
|
|
 |
16
|
|
 |
17
|
Dennis Fetterly , Mark Manasse , Marc Najork, Spam, damn spam, and statistics: using statistical analysis to locate spam web pages, Proceedings of the 7th International Workshop on the Web and Databases: colocated with ACM SIGMOD/PODS 2004, June 17-18, 2004, Paris, France
[doi> 10.1145/1017074.1017077]
|
 |
18
|
|
| |
19
|
|
| |
20
|
|
| |
21
|
Gyöngyi, Z. and Garcia-Molina, H. 2005b. Web spam taxonomy. In Proceedings of the 1st International Workshop on Adversarial Information Retrieval on the Web (AIRWeb'05).
|
| |
22
|
|
| |
23
|
Henzinger, M., Motwani, R., and Silverstein, C. 2003. Challenges in Web search engines. In Proceedings of the 18th International Joint Conference on Artificial Intelligence (IJCAI'03). 1573--1579.
|
| |
24
|
Karp, R. M. 1972. Reducibility among combinatorial problems. In Complexity of Computer Computations.
|
 |
25
|
|
| |
26
|
Langville, A. and Meyer, C. 2004. Deeper inside pagerank. Intern. Math. 1, 3, 335--380.
|
 |
27
|
|
| |
28
|
Page, L., Brin, S., Motwani, R., and Winograd, T. 1998. The pagerank citation ranking: Bringing order to the Web. Tech. rep., Stanford University.
|
| |
29
|
Scott, J. 2000. Social Network Analysis Handbook. Sage Publications Inc.
|
| |
30
|
Thompson, A. C. 1996. Minkowski Geometry. Cambridge University Press, Cambridge, UK.
|
| |
31
|
Wasserman, S. and Faust, K. 1994. Social Network Analysis. Cambridge University Press, Cambridge, UK.
|
 |
32
|
|
 |
33
|
|
| |
34
|
Zhang, H., Goel, A., Govindan, R., Mason, K., and Roy, B. V. 2004. Making eigenvector-based reputation systems robust to collusion. In Proceedings of the 3rd Workshop on Algorithms and Models for the Web Graph (WAW'04). Lecture Notes in Computer Science, vol. 3243. Springer, 92--104.
|
|