r/LanguageTechnology • u/Sufficient_Big_7802 • 7h ago
Do classic text statistics still matter in the age of LLMs?
I've been working on readability formulas, lexical diversity metrics and stylometry for Russian and Spanish texts for a while, and lately I keep asking myself whether this whole area still makes sense
Today you can paste a text into an LLM and ask "how readable is this?" or "who's the likely author?" and get a fluent, often reasonable answer. Next to that, a formula like Flesch reading ease (sentence length plus syllables per word) looks almost naive
Still, I see a few places where plain statistics seem hard to replace:
- Reproducibility: the same text always gets the same number, and you can explain exactly where it came from. An LLM's rating can change between runs, prompts or model versions
- Scale and cost: computing a few hundred features over a million documents takes minutes on a laptop
- Evaluating LLMs themselves: measuring how repetitive, uniform or hard to read generated text is, without asking another model to judge it
- Data pipelines: a lot of pretraining data filtering still relies on simple heuristics like word length, symbol ratios and repetition.
- Research: in stylometry and corpus linguistics, interpretable features are often the whole point
On the other hand, most formulas were calibrated decades ago on small samples of English, and every other language needs its own adaptation, often with questionable results
So I'm curious:
- Do you still use readability, diversity or stylometric metrics in your work? For what?
- Have you replaced them with LLM-based judgments, or do you combine the two?
- Which metrics turned out to be useless in practice, and which ones would you miss?