ACM Home Page
Please provide us with feedback. Feedback
Meme-tracking and the dynamics of the news cycle
Full text MovMov (18:51),  PdfPdf (668 KB)
Source
International Conference on Knowledge Discovery and Data Mining archive
Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining table of contents
Paris, France
SESSION: Research track papers table of contents
Pages 497-506  
Year of Publication: 2009
ISBN:978-1-60558-495-9
Authors
Jure Leskovec  Stanford University, Stanford, CA, USA
Lars Backstrom  Cornell University, Ithaca, NY, USA
Jon Kleinberg  Cornell University, Ithaca, NY, USA
Sponsors
ACM: Association for Computing Machinery
SIGKDD: ACM Special Interest Group on Knowledge Discovery in Data
SIGMOD: ACM Special Interest Group on Management of Data
Publisher
ACM  New York, NY, USA
Bibliometrics
Downloads (6 Weeks): 91,   Downloads (12 Months): 230,   Citation Count: 0
Additional Information:

abstract   references   index terms   collaborative colleagues  

Tools and Actions: Request Permissions Request Permissions    Review this Article  
DOI Bookmark: Use this link to bookmark this Article: http://doi.acm.org/10.1145/1557019.1557077
What is a DOI?

ABSTRACT

Tracking new topics, ideas, and "memes" across the Web has been an issue of considerable interest. Recent work has developed methods for tracking topic shifts over long time scales, as well as abrupt spikes in the appearance of particular named entities. However, these approaches are less well suited to the identification of content that spreads widely and then fades over time scales on the order of days - the time scale at which we perceive news and events.

We develop a framework for tracking short, distinctive phrases that travel relatively intact through on-line text; developing scalable algorithms for clustering textual variants of such phrases, we identify a broad class of memes that exhibit wide spread and rich variation on a daily basis. As our principal domain of study, we show how such a meme-tracking approach can provide a coherent representation of the news cycle - the daily rhythms in the news media that have long been the subject of qualitative interpretation but have never been captured accurately enough to permit actual quantitative analysis. We tracked 1.6 million mainstream media sites and blogs over a period of three months with the total of 90 million articles and we find a set of novel and persistent temporal patterns in the news cycle. In particular, we observe a typical lag of 2.5 hours between the peaks of attention to a phrase in the news media and in blogs respectively, with divergent behavior around the overall peak and a "heartbeat"-like pattern in the handoff between news and blogs. We also develop and analyze a mathematical model for the kinds of temporal variation that the system exhibits.


REFERENCES

Note: OCR errors may be found in this Reference List extracted from the full text article. ACM has opted to expose the complete List rather than only correct and linked references.

 
1
Supporting website: http://memetracker.org
2
 
3
E. Adar, L. Zhang, L. Adamic, R. Lukose. Implicit structure and dynamics of blogspace. Wks. Weblogging Ecosystem'04.
 
4
R. Albert and A.-L. Barabási. Statistical mechanics of complex networks. Rev. of Modern Phys., 74:47--97, 2002.
 
5
J. Allan (ed). Topic Detection and Tracking. Kluwer, 2002.
 
6
L. Bennett. News: The Politics of Illusion. A. B. Longman (Classics in Political Science), seventh edition, 2006.
7
 
8
 
9
 
10
11
 
12
M. Gamon, S. Basu, D. Belenko, D. Fisher, M. Hurst, and A. C. Kanig. Blews: Using blogs to provide context for news articles. In ICWSM '08, 2008.
 
13
N. Godbole, M. Srinivasaiah, and S. Skiena. Large-scale sentiment analysis for news and blogs. In ICWSM '07, 2007.
14
 
15
J. Harsin. The rumour bomb: Theorising the convergence of new and old trendsin mediated U.S. politics. Southern Review: Communication, Politics and Culture,39(2006).
 
16
17
 
18
M. Kot. Elements of Mathematical Ecology. Cambridge University Press, 2001.
 
19
B. Kovach and T. Rosenstiel. Warp Speed: America in the Age of Mixed Media. Century Foundation Press, 1999.
20
 
21
M. Lacker and C. Peskin. Control of ovulation number in a model of ovarian follicularmaturation. In AMS Symposium on Mathematical Biology,pages 21--32, 1981.
 
22
P.F. Lazarsfeld, B. Berelson, and H. Gaudet. The People's Choice. Duell, Sloan, and Pearce, 1944.
 
23
J. Leskovec, M. McGlohon, C. Faloutsos, N. Glance, M. Hurst. Cascading behavior in large blog graphs. SDM'07.
 
24
R. D. Malmgren, D. B. Stouffer, A. Motter, and L. A. N. Amaral. A poissonian explanation for heavy tails in e-mail communication. PNAS, to appear, 2008.
 
25
J. Schmidt. Blogging practices: An analytical framework. Journal of Computer-Mediated Communication, 12(4), 2007.
 
26
J. Singer. The political j-blogger. Journalism, 6(2005).
 
27
Spinn3r API. http://www.spinn3r.com. 2008.
 
28
M. L. Stein, S. Paterno, and R. C. Burnett. Newswriter's Handbook: An Introduction to Journalism. Blackwell, 2006.
 
29
A. Vazquez, J. G. Oliveira, Z. Deszo, K.-I. Goh, I. Kondor, and A.-L. Barabasi. Modeling bursts and heavy tails in human dynamics. Physical Review E, 73(036127), 2006.
30
31
 
32
F. Wu and B. Huberman. Novelty and collective attention. Proc. Natl. Acad. Sci. USA, 104, 2007.

Collaborative Colleagues:
Jure Leskovec: colleagues
Lars Backstrom: colleagues
Jon Kleinberg: colleagues