|
ABSTRACT
Our research examines a predictive machine learning approach for financial news articles analysis using several different textual representations: bag of words, noun phrases, and named entities. Through this approach, we investigated 9,211 financial news articles and 10,259,042 stock quotes covering the S&P 500 stocks during a five week period. We applied our analysis to estimate a discrete stock price twenty minutes after a news article was released. Using a support vector machine (SVM) derivative specially tailored for discrete numeric prediction and models containing different stock-specific variables, we show that the model containing both article terms and stock price at the time of article release had the best performance in closeness to the actual future stock price (MSE 0.04261), the same direction of price movement as the future price (57.1% directional accuracy) and the highest return using a simulated trading engine (2.06% return). We further investigated the different textual representations and found that a Proper Noun scheme performs better than the de facto standard of Bag of Words in all three metrics.
REFERENCES
Note: OCR errors may be found in this Reference List extracted from the full text article. ACM has opted to expose the complete List rather than only correct and linked references.
| |
1
|
Bishop, C. M. and Tipping, M. E. 2003. Bayesian Regression and Classification. IOS Press, Amsterdam.
|
| |
2
|
Burns, D. and Wutkowski, K. Nov. 15, 2005. Schwab to miss forecast, fined by NYSE. http://biz.yahoo.com/rb/051115/financial_schwab.html?.v=3.
|
| |
3
|
Cho, V. 1999. Knowledge Discovery from Distributed and Textual Data. Tech. rep. Department of Computer Science. Hong Kong University of Science and Technology.
|
| |
4
|
Cho, V., Wuthrich, B., and Zhang, J. 1998. Text processing for classification. J. Computat. Intel. Fin. 26.
|
 |
5
|
|
| |
6
|
Fama, E. 1964. The behavior of stock market prices. Tech. rep. Graduate School of Business, University of Chicago.
|
| |
7
|
|
| |
8
|
|
| |
9
|
Gidofalvi, G. 2001. Using news articles to predict stock price movements. Tech rep. Department of Computer Science and Engineering, University of California, San Diego.
|
| |
10
|
|
| |
11
|
Antonina Kloptchenko , Tomas Eklund , Jonas Karlsson , Barbro Back , Hannu Vanharanta , Ari Visa, Combining data and text mining techniques for analysing financial reports: Research Articles, International Journal of Intelligent Systems in Accounting and Finance Management, v.12 n.1, p.29-41, January 2004
[doi> 10.1002/isaf.v12:1]
|
| |
12
|
Lavrenko, V., Schmill, M., Lawrie, D., and Ogilvie, P. 2000b. Mining of concurrent text and time series. In Proceedings of the 6th ACM International Conference on Knowledge Discovery and Data Mining (KDD).
|
 |
13
|
Victor Lavrenko , Matt Schmill , Dawn Lawrie , Paul Ogilvie , David Jensen , James Allan, Language models for financial news recommendation, Proceedings of the ninth international conference on Information and knowledge management, p.389-396, November 06-11, 2000, McLean, Virginia, United States
[doi> 10.1145/354756.354845]
|
| |
14
|
Le Moigno, S., Charlet, J., Bourigualt, D., Degoulet, P., and Jaulent, M.-C. 2002. Terminology extraction from text to build an ontology in surgical intensive care. In Proceedings of the AMIA Symposium.
|
| |
15
|
LeBaron, B., Arthur, W. B., and Palmer, R. 1999. Time series properties of an artificial stock market. J. Econ. Dynam. Contr. 23, 9--10, 1487--1516.
|
| |
16
|
Malkiel, B. G. 1973. A Random Walk Down Wall Street. W.W. Norton, New York.
|
| |
17
|
McDonald, D. M., Chen, H., and Schumaker, R. P. 2005. Transforming open-source documents to terror networks: The Arizona TerrorNet. In Proceedings of the American Association for Artificial Intelligence Conference Spring Symposia.
|
| |
18
|
|
 |
19
|
|
| |
20
|
Pai, P.-F. and Lin, C.-S. 2005. A hybrid ARIMA and support vector machines model in stock price forecasting. Omega 33, 6, 497--505.
|
| |
21
|
|
| |
22
|
Sekine, S. and Nobata, C. 2003. Definition, dictionaries and tagger for extended named entity hierarchy. In Proceedings of the International Conference on Language Resources and Evaluation.
|
| |
23
|
Seo, Y.-W., Giampapa, J., and Sycara, K. 2002. Text classification for intelligent portfolio management. Tech rep. Robotics Institute, Carnegie Mellon University.
|
| |
24
|
Tay, F. and Cao, L. 2001. Application of support vector machines in financial time series forecasting. Omega 29, 309--317.
|
| |
25
|
Technical-Analysis. 2005. The Trader's Glossary of Technical Terms and Topics. http://www.traders.com/documentation/RESource_docs/glossary/glossary.html.
|
| |
26
|
Thomas, J. D. and Sycara, K. 2002. Integrating genetic algorithms and text learning for financial prediction. In Proceedings of the Genetic and Evolutionary Computation Conference (GECCO).
|
| |
27
|
|
| |
28
|
Vanschoenwinkel, B. 2003. A discrete kernel approach to support vector machine learning in language independent named entity recognition. Tech. rep. Computational Modeling Lab, Vrije Universiteit, Brussels.
|
|