About DeepFij
DeepFij is college basketball scores, schedules, power ratings, and machine-learning game predictions for every Division I program, updated daily through the season.
This project has been a labor of love for 25 years. Portions of what you see on the site have been implemented partially in one form or another in C, C++, Java, Scala, Ruby, Python, and Node. The interface has been X/Motif, Java Swing, a Java webapp, a Scala/Play webapp, Angular and React. The data store has been flat text files, MySQL, MongoDB, and Postgres.
For a while the project existed to me simply as a known problem, one with which I could experiment to learn a new technology. To some extent, that's how it remains. But now the new technology is LLM's. Thanks to their magic it is now in a state suitable for for others to see.
Be advised it's not my day job. It could disappear any day, and is always presented on a best efforts basis. No warranty expressed or implied.
Almost nothing on this site is terribly novel. (If anything is, I think maybe the game chart). The models listed as `Massey` and `Bradley-Terry` have been 'discovered' independently many times by people who were interested by the problem of ranking and estimating by paired comparisons, a group which included me. The names, and citations I include below, come from the first or best published reference I know. The pace, efficiency, and all the other modern stats are due to work by Dean Oliver and Ken Pomeroy. Those guys, along with Keith Massey, are the OGs of college basketball ratings.
Predictions come from a family of models — classical Massey and Bradley-Terry ratings alongside gradient-boosted ML models — and every pre-game prediction is scored against the final result and the closing line, in public, on the model performance page.
Game data is sourced from ESPN's public APIs. DeepFij is an independent analytics project and is not affiliated with ESPN or the NCAA. This is not betting advice.
Relevant Literature
If you feel these references are incomplete or I fail to give proper credit, please let me know at deepfij@gmail.com.
Team ratings
- Bradley, R. A. & Terry, M. E. (1952). “Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons.” Biometrika 39, 324–345. The paired-comparison model behind our Bradley‑Terry ratings.
- Harville, D. (1977). “The Use of Linear-Model Methodology to Rate High School or College Football Teams.” Journal of the American Statistical Association 72, 278–289. The academic foundation for least-squares team ratings with home advantage.
- Massey, K. (1997). Statistical Models Applied to the Rating of Sports Teams. Honors thesis, Bluefield College. The least-squares rating system implemented as our Massey model.
- Stern, H. (1991). “On the Probability of Winning a Football Game.” The American Statistician 45, 179–183. Margins of victory are roughly normal — how a predicted margin becomes a win probability.
Pace & efficiency
- Oliver, D. (2004). Basketball on Paper. Brassey's. The possession framework and the four factors.
- Kubatko, J., Oliver, D., Pelton, K. & Rosenbaum, D. T. (2007). “A Starting Point for Analyzing Basketball Statistics.” Journal of Quantitative Analysis in Sports 3(3). The possession estimator and per-possession efficiency conventions we use.
- Hoerl, A. E. & Kennard, R. W. (1970). “Ridge Regression: Biased Estimation for Nonorthogonal Problems.” Technometrics 12, 55–67. The regularization behind our opponent-adjusted efficiency ratings.
Machine learning & evaluation
- Chen, T. & Guestrin, C. (2016). “XGBoost: A Scalable Tree Boosting System.” Proceedings of KDD '16. The gradient-boosted trees powering our ML predictions.
- Akiba, T., Sano, S., Yanase, T., Ohta, T. & Koyama, M. (2019). “Optuna: A Next-generation Hyperparameter Optimization Framework.” Proceedings of KDD '19. Hyperparameter tuning for model training.
- Brier, G. W. (1950). “Verification of Forecasts Expressed in Terms of Probability.” Monthly Weather Review 78, 1–3. The Brier scores on the model performance page.
- Gneiting, T. & Raftery, A. E. (2007). “Strictly Proper Scoring Rules, Prediction, and Estimation.” Journal of the American Statistical Association 102, 359–378. Why we score probability forecasts the way we do.
Tempo-adjusted efficiency ratings were popularized in college basketball by Ken Pomeroy (kenpom.com) and Bart Torvik (barttorvik.com), whose public work shaped the conventions this whole genre of analysis follows.