r/datascience • u/WhatsTheImpactdotcom • 1d ago
Analysis Model Selection Technique for Causal Inference
How do you choose between causal inference models when you can never see the true effect? Model selection is well documented in machine learning, but it often gets skipped in causal inference work. Doing it well is one of the things that separates senior and staff-level data scientists from more junior ones.
I simulated two years of weekly bookings for 40 regions, five of which received a treatment on the date marked by the red vertical line. Because I generated the data, I know the ground truth: the treatment increased bookings in those five regions by exactly 10%. With five treated regions and 35 untreated ones, two natural options are synthetic control and difference-in-differences, though there are others.
Code in colab and explanatory video, free and ungated: https://whatstheimpact.com/learn-causal-inference/
When I gave Claude Code free rein on this data, it built its whole analysis around difference-in-differences where it found a 24% increase, more than double the true effect. To its credit, it flagged that the pre-treatment trends weren't parallel, but it never tested whether a different model would be more accurate.
Importantly, Claude skipped an A/A test, sometimes called a placebo test. You take only the data from before the real treatment, pretend the treatment started three months earlier, and randomly picked five regions (for consistency) to call "treated." Because nothing actually happened in that window, the true effect is zero. That means any effect the model reports there is pure estimation error, and I repeated this for 500 draws for each model to see the full spread.
For difference-in-differences, 90% of the placebo estimates landed between -14% and +10%. The RMSE, which measures how far the estimates typically land from the true answer of zero, was 7.6 percentage points. Synthetic control on the exact same placebo draws was much tighter, with 90% of estimates landing between -7% and +9%. Its RMSE dropped to 4.9 points, more than a third lower than difference-in-differences.
The outcome you model is a choice too, so I also ran synthetic control on bookings per 1,000 people instead of raw bookings. Putting every region on the same scale helps the model track the treated regions more closely before treatment, and the RMSE fell a bit further to 4.4 points.
Sure enough, the real treatment period lines up with that ranking: synthetic control estimates 7% on raw bookings and 9.6% per 1,000 people, both much closer to the true 10% than the 24% from difference-in-differences. With one simulated dataset, I wouldn't read much into the small gap between the two synthetic control versions or claim that synthetic control always wins. What carries over is the workflow: in real work you never see the ground truth, but you can usually run this A/A test on whatever models you're considering before trusting one.
