| An unsupervised framework for extracting and normalizing product attributes from multiple web sites |
| Full text |
Pdf
(332 KB)
|
Source
|
Annual ACM Conference on Research and Development in Information Retrieval
archive
Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval
table of contents
Singapore, Singapore
SESSION: Web search--1
table of contents
Pages 35-42
Year of Publication: 2008
ISBN:978-1-60558-164-4
|
|
Authors
|
|
Tak-Lam Wong
|
The Chinese University of Hong Kong, Hong Kong, Hong Kong
|
|
Wai Lam
|
The Chinese University of Hong Kong, Hong Kong, Hong Kong
|
|
Tik-Shun Wong
|
The Chinese University of Hong Kong, Hong Kong, Hong Kong
|
|
| Sponsors |
|
| Publisher |
|
| Bibliometrics |
Downloads (6 Weeks): 55, Downloads (12 Months): 553, Citation Count: 0
|
|
|
ABSTRACT
We have developed an unsupervised framework for simultaneously extracting and normalizing attributes of products from multiple Web pages originated from different sites. Our framework is designed based on a probabilistic graphical model that can model the page-independent content information and the page-dependent layout information of the text fragments in Web pages. One characteristic of our framework is that previously unseen attributes can be discovered from the clue contained in the layout format of the text fragments. Our framework tackles both extraction and normalization tasks by jointly considering the relationship between the content and layout information. Dirichlet process prior is employed leading to another advantage that the number of discovered product attributes is unlimited. An unsupervised inference algorithm based on variational method is presented. The semantics of the normalized attributes can be visualized by examining the term weights in the model. Our framework can be applied to a wide range of Web mining applications such as product matching and retrieval. We have conducted extensive experiments from four different domains consisting of over 300 Web pages from over 150 different Web sites, demonstrating the robustness and effectiveness of our framework.
REFERENCES
Note: OCR errors may be found in this Reference List extracted from the full text article. ACM has opted to expose the complete List rather than only correct and linked references.
| |
1
|
I. Bhattacharya and L. Getoor. A latent dirichlet model for unsupervised entity resolution. In Proceedings of the 2006 SIAM International Conference on Data Mining, pages 47--58, 2006.
|
 |
2
|
|
| |
3
|
D. Blei and M. Jordan. Variational inference for dirichlet process mixtures. Bayesian Analysis, 1(1):121--144, 2006.
|
| |
4
|
|
| |
5
|
|
| |
6
|
J. Ishwaran and L. James. Gibbs sampling methods for stick-breaking priors. Journal of the American Statistical Association, 96(453):161--174, 2001.
|
| |
7
|
|
| |
8
|
A. McCallum and D. Jensen. A note on the unification of information extraction and data mining using conditional-probability, relational models. In Proceedings of the IJCAI Workshop on Learning Statistical Models from Relational Data, 2003.
|
| |
9
|
|
| |
10
|
K. Probst, M. K. R. Ghai, A. Fano, and Y. Liu. Semi-supervised learning of attribute-value pairs from product descriptions. In Proceedings of the Twentieth International Joint Conference on Artificial Intelligence, pages 2838--2843, 2007.
|
 |
11
|
|
| |
12
|
S. Sarawagi and W. Cohen. Semi-markov conditional random fields for information extraction. In Advances in Neural Information Processing Systems 17, Neural Information Processing Systems, 2004.
|
| |
13
|
|
 |
14
|
Charles Sutton , Khashayar Rohanimanesh , Andrew McCallum, Dynamic conditional random fields: factorized probabilistic models for labeling and segmenting sequence data, Proceedings of the twenty-first international conference on Machine learning, p.99, July 04-08, 2004, Banff, Alberta, Canada
[doi> 10.1145/1015330.1015422]
|
| |
15
|
Y. Teh, M. Jordan, M. Beal, and D. Blei. Hierarchical dirichlet processes. Journal of the American Statistical Association, 101:1566--1581, 2006.
|
| |
16
|
Ben Wellner , Andrew McCallum , Fuchun Peng , Michael Hay, An integrated, conditional model of information extraction and coreference with application to citation matching, Proceedings of the 20th conference on Uncertainty in artificial intelligence, p.593-601, July 07-11, 2004, Banff, Canada
|
 |
17
|
|
 |
18
|
|
 |
19
|
Jun Zhu , Bo Zhang , Zaiqing Nie , Ji-Rong Wen , Hsiao-Wuen Hon, Webpage understanding: an integrated approach, Proceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining, August 12-15, 2007, San Jose, California, USA
[doi> 10.1145/1281192.1281288]
|
|