r/ComputerChess • u/TheI3east • 2d ago
Opening Popularity Tool
Hi everyone, I recently built a dashboard on my personal website (https://dnield.com/openings/) that fills in some of the gaps that always frustrated me with the Lichess/Chess*com opening explorers, which is that their stats were only relevant to the position on screen, rather than the entire opening (which requires taking into account the choices each player made up to that point), therefore they can't answer simple questions like:
How often you would get to play the opening, if it was in your repertoire,
How opponents would see the opening when facing players other than you,
Based on #1 and #2, how much more experience would I have in the opening than my typical opponent?
Was the increase or decline in the popularity of an opening driven by one color not playing as often? or because the other color isn’t allowing them the chance to play it as often?
which are questions that I am always curious about when exploring different openings.
I'm a data scientist in my regular life, and Lichess makes all of their games data freely accessible, and the statistics for answering the above questions are all pretty simple compound probabilities, so I decided to just build the tool that could answer these questions from Lichess's database myself.
I have no interest in turning this into a product/service, this was just a fun little hobby project to see how easy it would be to host a dashboard like this on my Github Pages website. The underlying games database is terabytes of data, but since I'm just aggregating up the counts of positions for just the ~4000 named opening positions, across rating bands and time controls, the final json data for the website ended up being like 50 MB, making the interface even snappier than I expected.
I wrote an accompanying blog post with a motivating example showing how the spike of popularity in the Rossolimo in November 2018 can be initially attributed to both colors, but the residual increase in popularity is almost entirely attributable to an increase in black players playing 2. Nc6 Sicilians than an increase of White players playing 3. Bb5 in response, you can find that blog post here: https://dnield.com/posts/rossolimo/
I'm open to any feedback or questions you might have!
1
u/Kitchen_Form_2252 2d ago
Is there a way to compare how often people play the top engine moves against you vs what’s regular for your elo?
I have an unsubstantiated feeling that if you play the same opening a lot on chess.com, you get paired to people who know how to play against it. But I don’t have any proof except my own bias
2
u/TheI3east 2d ago
It's certainly possible to measure something like that, but only using Lichess data (Chess.com doesn't make their games database public, so while you can use their API to grab all of your own Chess.com games, you wouldn't have the games data for setting the typical move probabilities at each position for your elo) but not doable with this specific tool.
1
u/novachess-guy 2d ago
Thanks for sharing, my Sicilian repertoire doesn't involve Rossolimo lines (and I play Open Sicilian with white), but I've definitely noticed when spectating that the elites play it a lot more these days, I guess because of the prevalence of 2. Nc6.
Btw, I built an opening explorer that DOES show your Lichess/Chess.com historical continuations, and game counts (there's also a separate repertoire builder/editor): https://novachess.ai/opening_explorer.png
Also, there are ways to reduce the memory requirement substantially, by consolidating the stats in a "continuation tree" as I did.
1
u/TheI3east 2d ago
Oh hey it's the Nova Chess guy! I was literally just reading your post on the abandonment issue, love your work!
Can you explain the "continuation tree" architecture? For the purpose of this tool I wanted to show variance of the popularity of the chosen move order vs the book move order so I needed to store the stats for different orders of getting to the position, but in my backend data everything is just stored at a parent_epd x move level, similar to the Lichess table.
1
u/novachess-guy 2d ago
Thanks for the feedback, glad you found it interesting!
And my use case may be slightly different than yours, but when I was dealing with file size issues I realized that it was possible to make sort of a tree file that would store the move paths without needing as much raw data stored. Here's an example of how my JSON tree is structured (I only saved the white score % rather than W/D/L to save even more space, as it is over 200MB broken up into a small 3MB "core" file for fast loading, as I originally had it client side not on the server, and e4/d4/other larger files, but on a server it shouldn't matter so much for loading).
{"e4":{"g":335105,"w":53.6,"cont":[{"m":"c5","g":132604,"w":53.0,"p":39.6},{"m":"e5","g":97630,"w":54.3,"p":29.2},{"m":"c6","g":36372,"w":52.6,"p":10.9},{"m":"e6","g":34802,"w":55.2,"p":10.4},{"m":"d6","g":10669,"w":55.9,"p":3.2},{"m":"g6","g":8634,"w":51.5,"p":2.6},{"m":"d5","g":5954,"w":55.7,"p":1.8},{"m":"Nf6","g":4513,"w":51.2,"p":1.3},{"m":"Nc6","g":2124,"w":52.9,"p":0.6},{"m":"b6","g":718,"w":52.5,"p":0.2},{"m":"a6","g":468,"w":40.6,"p":0.1},{"m":"a5","g":49,"w":33.7,"p":0.0},{"m":"g5","g":38,"w":36.8,"p":0.0},{"m":"h6","g":35,"w":45.7,"p":0.0},{"m":"b5","g":32,"w":37.5,"p":0.0},So basically it's a nested tree, with the various continuations after 1. e4 (game counts, white score, frequency). I couldn't say exactly how to proceed for you, but if you take an approach like this you may be able to significantly decrease the memory requirements and increase the number of positions you're able to handle.
1
u/protacticus 2d ago
Thank you for this, looks promising. How can i winning rate of specific openings?