Projects and research code related to statistics, econometrics, and machine learning. The repositories vary in scope and implementation language; many explore a specific estimator, design choice, or computational method.
📐 Econometrics and causal inference: LATE, late_iv, tworeg, fuzzy, gsynth, smooth-operator, and simcheck. These repositories study identification, sensitivity, treatment-effect estimation, smoothing, and simulation-based checks.
🎯 Calibration, scoring, and decision rules: calibre, streamcal, fairlex, rank-preserving-calibration, score_signs, winference, optimal-classification-cutoffs, optimal_cuts, and queue-shift. They address probability calibration, multiclass thresholds, pairwise rankings, fairness constraints, and deployment decisions.
🌲 Stable and robust machine learning: stable-cart, robust-cart, stableboost, bcr, stable-gen, dct, act, mpsam, sam-lasso, treegptq, stagecoachml, and ensemble-proximity. The common question is when resampling, consistency training, or constrained updates make fitted models less sensitive to the sample.
🔎 Matching, joins, and dimension reduction: preclink, setjoin, onetomany, pyppann, pyppur, incline, lookahead-cart, lookahead-kmeans, hbw, and alsgls. These tools make linkage objectives, projection-pursuit reductions, nonparametric smoothing, and structured search explicit.
🧪 Design, measurement, and data collection: fewlab, fewest_domains, optimal_data_collection, prop_male, hybrid, and lowdimtraining. They examine how sampling, labeling, training updates, and stopping rules affect precision and evidence value.
📚 Replications, benchmarks, and teaching materials: econometric_bench, benchmarking-benchmarks, guess, dann, total_error, bagged_fsr, bagged_mp, deliberately, and ds. These projects preserve runnable examples, reproduce published or proposed calculations, and test methods under controlled finite samples.
🤖 R interfaces: rmcp exposes selected R capabilities through an MCP server.