|
ABSTRACT
Tracking new topics, ideas, and "memes" across the Web has been an issue of considerable interest. Recent work has developed methods for tracking topic shifts over long time scales, as well as abrupt spikes in the appearance of particular named entities. However, these approaches are less well suited to the identification of content that spreads widely and then fades over time scales on the order of days - the time scale at which we perceive news and events. We develop a framework for tracking short, distinctive phrases that travel relatively intact through on-line text; developing scalable algorithms for clustering textual variants of such phrases, we identify a broad class of memes that exhibit wide spread and rich variation on a daily basis. As our principal domain of study, we show how such a meme-tracking approach can provide a coherent representation of the news cycle - the daily rhythms in the news media that have long been the subject of qualitative interpretation but have never been captured accurately enough to permit actual quantitative analysis. We tracked 1.6 million mainstream media sites and blogs over a period of three months with the total of 90 million articles and we find a set of novel and persistent temporal patterns in the news cycle. In particular, we observe a typical lag of 2.5 hours between the peaks of attention to a phrase in the news media and in blogs respectively, with divergent behavior around the overall peak and a "heartbeat"-like pattern in the handoff between news and blogs. We also develop and analyze a mathematical model for the kinds of temporal variation that the system exhibits.
REFERENCES
Note: OCR errors may be found in this Reference List extracted from the full text article. ACM has opted to expose the complete List rather than only correct and linked references.
| |
1
|
Supporting website: http://memetracker.org
|
 |
2
|
|
| |
3
|
E. Adar, L. Zhang, L. Adamic, R. Lukose. Implicit structure and dynamics of blogspace. Wks. Weblogging Ecosystem'04.
|
| |
4
|
R. Albert and A.-L. Barabási. Statistical mechanics of complex networks. Rev. of Modern Phys., 74:47--97, 2002.
|
| |
5
|
J. Allan (ed). Topic Detection and Tracking. Kluwer, 2002.
|
| |
6
|
L. Bennett. News: The Politics of Illusion. A. B. Longman (Classics in Political Science), seventh edition, 2006.
|
 |
7
|
|
| |
8
|
|
| |
9
|
|
| |
10
|
|
 |
11
|
|
| |
12
|
M. Gamon, S. Basu, D. Belenko, D. Fisher, M. Hurst, and A. C. Kanig. Blews: Using blogs to provide context for news articles. In ICWSM '08, 2008.
|
| |
13
|
N. Godbole, M. Srinivasaiah, and S. Skiena. Large-scale sentiment analysis for news and blogs. In ICWSM '07, 2007.
|
 |
14
|
Daniel Gruhl , R. Guha , David Liben-Nowell , Andrew Tomkins, Information diffusion through blogspace, Proceedings of the 13th international conference on World Wide Web, May 17-20, 2004, New York, NY, USA
[doi> 10.1145/988672.988739]
|
| |
15
|
J. Harsin. The rumour bomb: Theorising the convergence of new and old trendsin mediated U.S. politics. Southern Review: Communication, Politics and Culture,39(2006).
|
| |
16
|
|
 |
17
|
|
| |
18
|
M. Kot. Elements of Mathematical Ecology. Cambridge University Press, 2001.
|
| |
19
|
B. Kovach and T. Rosenstiel. Warp Speed: America in the Age of Mixed Media. Century Foundation Press, 1999.
|
 |
20
|
|
| |
21
|
M. Lacker and C. Peskin. Control of ovulation number in a model of ovarian follicularmaturation. In AMS Symposium on Mathematical Biology,pages 21--32, 1981.
|
| |
22
|
P.F. Lazarsfeld, B. Berelson, and H. Gaudet. The People's Choice. Duell, Sloan, and Pearce, 1944.
|
| |
23
|
J. Leskovec, M. McGlohon, C. Faloutsos, N. Glance, M. Hurst. Cascading behavior in large blog graphs. SDM'07.
|
| |
24
|
R. D. Malmgren, D. B. Stouffer, A. Motter, and L. A. N. Amaral. A poissonian explanation for heavy tails in e-mail communication. PNAS, to appear, 2008.
|
| |
25
|
J. Schmidt. Blogging practices: An analytical framework. Journal of Computer-Mediated Communication, 12(4), 2007.
|
| |
26
|
J. Singer. The political j-blogger. Journalism, 6(2005).
|
| |
27
|
Spinn3r API. http://www.spinn3r.com. 2008.
|
| |
28
|
M. L. Stein, S. Paterno, and R. C. Burnett. Newswriter's Handbook: An Introduction to Journalism. Blackwell, 2006.
|
| |
29
|
A. Vazquez, J. G. Oliveira, Z. Deszo, K.-I. Goh, I. Kondor, and A.-L. Barabasi. Modeling bursts and heavy tails in human dynamics. Physical Review E, 73(036127), 2006.
|
 |
30
|
|
 |
31
|
Xuanhui Wang , ChengXiang Zhai , Xiao Hu , Richard Sproat, Mining correlated bursty topic patterns from coordinated text streams, Proceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining, August 12-15, 2007, San Jose, California, USA
[doi> 10.1145/1281192.1281276]
|
| |
32
|
F. Wu and B. Huberman. Novelty and collective attention. Proc. Natl. Acad. Sci. USA, 104, 2007.
|
|