Notes · markets · learning from data

Quantitative finance

A personal reading list and research notebook on statistical learning, trading, portfolio construction, and market microstructure.

A personal survey

Ideas for finding signal in noisy markets.

Quantitative finance is a fascinating area for anyone who has the mathematical background to understand quant theory and has funds to invest. I decided to share some interesting papers, books, and websites here. These references cover a wide range of areas including machine learning, trading, portfolio optimization, and market microstructure. Most references are freely available on the web.

This is a personal survey from the point of someone with a good understanding of probability theory, statistics, signal processing, and related fields. The emphasis is on estimating future returns of financial instruments based on past price information, relative-value strategies, and other signals that are combined via state-of-the-art machine learning methods.

02 / Reading list

Books, papers, and resources

12 references
01

The Elements of Statistical Learning

T. Hastie, R. Tibshirani, J. H. Friedman · July 2003

This book can be considered the bible of quantitative finance for anyone looking for predictable patterns in financial markets. Key notions of overlearning, overfitting, and generalization ability are discussed, alongside a large collection of practical algorithms. It is a valuable reference and a good starting point.

What’s the catch? The area of machine learning and pattern recognition is huge. This book is just a starting point; you need to read the original papers to understand what’s happening at a deeper level.

02

ECE 901: Statistical Learning Theory

Lecture notes by Rob Nowak

I would recommend this if you want to understand the theory of statistical learning. Theory and practice tend to differ when it comes to machine learning, as they do in many disciplines. Theory gives insight into algorithm design, yet the performance of most algorithms is demonstrated via simulations. In spirit, I tend to agree with Vladimir N. Vapnik: “There’s nothing more practical than a good theory.”

03

News and Trading Rules

J. D. Thomas · PhD thesis, Carnegie Mellon University · 2003

I believe that financial markets are by and large efficient, and any single signal has little predictive power. Statistical learners tend to operate in the low-SNR regime: signals are weakly correlated with future returns and misclassification errors are large.

This thesis gave me the valuable perspective that learning algorithms should be robust to noise. With daily or weekly data sets, the goal is not to find the best single, potentially complex predictor, but to find many simple predictors that are marginally powerful and, when combined, provide predictive power and robustness. My experience has been that an optimized neural network can have much worse generalization ability than an ensemble of equally weighted linear predictors.

The thesis also offers a good lesson in quant research: have a null hypothesis (H0) and compute p-values and z-scores against it to test predictive power. I tend to use a heteroscedastic random walk as H0 in my tests. With a large enough data set, cross-validation may be sufficient.

04

Stacked Regressions

L. Breiman · Machine Learning, Springer · 1996

Ensemble methods are the name of the game in statistical learning. You should check out papers on bagging, boosting, and related variants. This paper proposes combining multiple predictors linearly with non-negative weights, chosen via leave-one-out cross-validation. The method is intuitive and relatively easy to implement.

05

Stock Market Prices Do Not Follow Random Walks

A. W. Lo and C. MacKinlay · Review of Financial Studies, vol. 1, pp. 41–66 · 1988

This paper introduced the Variance Ratio (VR) test. Suppose you have a time series and want to understand its character: is it trending, mean-reverting, or uncorrelated like a random walk? The VR test gives an answer and an asymptotic formula for statistical significance.

When applied to stocks, I find it remarkable that the VR test can reveal this information, whereas observing the autocorrelation function directly barely shows any temporal correlation. See reference 8 below for an excellent introduction to VR with examples.

06

Marketsci Blog

marketsci.wordpress.com

I love this blog. Suppose you ran a VR test on the S&P 500 index time series and found that it was strongly mean-reverting over 5–15 days. How are you going to exploit this seeming inefficiency? The blog finds and back-tests a multitude of indicators that can help. The strategies tend to be contrarian, betting on mean reversion one way or another.

The indicators include daily follow-through, various forms of moving average interpreted in a contrarian manner, and RSI(2). The remarkable thing is that these indicators gave good and statistically significant performance across all time periods in the last decade examined. My back-testing shows that combining some of these indicators using ensemble techniques gives decent out-of-sample performance.

While academic papers debunk the utility of technical indicators, I find it staggering that the strategies outlined in this blog work. My take is that a good technical indicator should be based on short-term price information relative to sample size. Market rules change over time, so ensuring statistical significance in the recent past is key.

07

Pairs Trading: Quantitative Methods and Analysis

G. Vidyamurthy · Wiley Finance · 2004

One of the best quant and trading books I’ve read. The author gives an excellent overview of pairs-trading methodology in simple language for a reader with a signal-processing background, with a useful section on Arbitrage Pricing Theory.

This book inspired me to try principal component analysis on stock-price time series, or stock returns, to decompose them into uncorrelated components.

08

A Computational Methodology for Modelling the Dynamics of Statistical Arbitrage

Andrew Neil Burgess · PhD thesis, London Business School · 1999

This not-so-well-known thesis is an excellent reference on cointegration, or relative-value strategies, for stocks. It also introduces the Variance Ratio test with excellent visuals and examples.

Two or more time series are cointegrated if a linear combination of them is stationary while the series themselves are not. This means profit may be possible by betting on the direction of linear combinations of stock prices. Cointegration is a generalization of pairs trading: pairs trading bets on the relative moves of two time series, while cointegration bets on linear combinations of two or more.

A cointegration bet does not always have to be for convergence. As long as the statistics of the stationary time series are known, one can bet in either direction. I recommend checking Principal Component Analysis (PCA) and other blind source-separation techniques for finding linear combinations that yield stationarity.

09

The Kelly Criterion in Blackjack, Sports Betting, and the Stock Market

E. O. Thorp · The 10th International Conference on Gambling and Risk Taking · Montreal · 1997

The most basic prediction problem deals with estimating future returns for a single time series. The next step is to bet on relationships between two time series, through pairs trading, and then on several series, through cointegration.

Assuming you estimated the mean and covariance of future returns for multiple instruments, how do you allocate your funds? The Kelly criterion addresses this by maximizing the growth rate of capital. You need to know, or estimate, the joint distribution of future returns; the mean and covariance suffice up to a second-order approximation.

A not-so-well-known property of the Kelly criterion is that it maximizes the Sharpe ratio under an L2-norm constraint on portfolio weights. I mathematically proved this myself and haven’t seen it mentioned anywhere. Ping me if you want to learn more.

10

Quantitative Equity Portfolio Management

L. B. Chincarini and Daehwan Kim · McGraw-Hill Library of Investment and Finance · 2006

Excellent introduction to factor models. I particularly liked the sections on market anomalies and popular fundamental factors. It is much more readable than a standard reference in this area, such as Grinold and Kahn. The approach is not very useful if you lack quantitative fundamental data.

11

How Markets Slowly Digest Changes in Supply and Demand

J. P. Bouchaud, J. D. Farmer, F. Lillo

Market microstructure survey. A must-read if you want to understand the price-formation process in stock markets. The first author has many interesting articles.

03 / Applied work

Research projects

3 projects

Some research projects I’ve worked on, with pointers to the approach outlined in the reading list above.

01

S&P 500 index daily return estimation from past prices

  • Find price-based and seasonal signals correlated with tomorrow’s return using a linear base learner.
  • Distinguish “fake patterns” from real ones by computing p-values and z-scores against a heteroscedastic random walk.
  • Combine multiple estimators via ensemble methods.
02

Trading cointegrated large-cap ETFs

  • Use principal component analysis to decompose daily returns.
  • Estimate the return of each component optimally with a linear model.
  • Apply Kelly betting based on estimated return means and covariances.
03

K Nearest-Neighbor (kNN) classifier from daily bars

  • Branch-and-bound-based optimized C++ implementation for a kNN classifier.
  • Daily bars seem to have statistically significant predictive power for Russell 2000 small-cap stocks but not for large caps.
  • Predictive power is strongest in 2000–2004 and diminishes over time.

This page was written before I took a full-time job at an algorithmic trading company in Dec. 2009.