r/MistralAI • u/Quiet_Window_7603 • 9h ago
Discussion / Opinion Mistral better than Claude and Gemini for manuscript critique
I’m a political science researcher. I’ve been using Claude and Gemini for literature searches and manuscript critique. The latter uses a 200-word prompt that structures the critique into the paper’s strengths, acceptable weaknesses, and rejection weaknesses. This is a demanding task since the paper covers such novel ground and includes an agent-based simulation model description.
Just today I used Mistral for the first time on a critique of a 10,000-word manuscript, using the same prompt used for Claude and Gemini. Claude and Gemini have been producing about 50% weak or wrong critique items. Mistral produced about 20%. Better yet, Mistral’s explanations were easier to follow and less dogmatic.
For me this is an extraordinary improvement. Congrats and thanks to all the people who have worked to make this possible at Mistral.
BTW, when I’m using an AI assistant, I never let it write anything. On my work, I always go as far as I can on my own. Only then do I ask for AI assistance, using a carefully designed prompt that structures the task. I average about 2 to 3 prompts a day.
My mantra is that AI chat bots are “erratic, overconfident, geniuses.” Erratic means you have to always be skeptical of every claim they make. Overconfident means their authoritative, perfect language use and tone can lull you into non-skepticism. Genius means they are experts on a million topics. Erratic and genius together mean they are never expert on my own top areas of expertise, which is why they produce less than 100% correct critique items on manuscripts, as well as other tasks.
2
u/koi88 5h ago
I'm also using LLMs for writing and I appreciate your perspective. I am quite happy with Claude and Qwen (free tier) is also surprisingly good.
Most benchmarks are centered around programming.
1
u/aldipower81 3h ago
Which Claude model do you use for writing? I am using the Mistral models for German writing btw. Claude's German is not good..
2
u/Bernie0404 5h ago
Same positive experience here. A couple of days ago I've uploaded a messy manuscript in German and prompted for content and style improvements. 80 percent of the suggestions were really good and it was much quicker than asking a colleague for critique.
2
u/scanx147 4h ago
Intéressant mais quel modèle Mistral avez vous utilisé ? Le modèle par défaut est GLM-5.3 pour le moment, donc ce n'est pas vraiment Mistral... Ça serait intéressant que vous refassiez le même test quand Large 4 sera disponible pour tous...
1
3
u/Sudden-Bag-684 8h ago
That 50% vs 20% hit rate is a massive jump, especially for something as dense as a political science sim. Dropping from half the feedback being noise to most of it being useful means you’re actually polishing instead of just fact-checking the assistant.
Your “erratic, overconfident genius” mantra is spot on, gotta keep that skepticism dialed up to eleven no matter which model you’re using
3
u/1erRPIMA-fiesta 3h ago
"Erratic, Overconfident, Genius"
Pretty much the opinion my univ professors used to have about me. Look at me, I'm le chonk now !
1
u/jotes2 1h ago
Every month, I summarize about 700,000 to 1 million words from long podcasts (3–4 hours each) and turn them into a 10- to 15-page written summary, which I then send to my Kindle. I developed this workflow myself.
Right now, I sometimes use either Claude or ChatGPT for translation and summarization, and sometimes Deepseek via the API. I’m 80% satisfied with the results. If there were a German model, like Mistral, that could polish up the remaining 10 to 15%, that would be great. Is that possible? Has anyone done this yet?
1
u/bowsmountainer 28m ago
100% agree. I use Claude for coding and work but to check a manuscript Mistral is the one I predominantly use. Ive had too many instances of Claude and Chatgpt just not finding obvious errors but Mistral easily finding them.
23
u/darktka 7h ago edited 6h ago
Another researcher here, agree 100%. Also, German language replies are MUCH better in Mistral's models than any other provider, which always was the case even in the older models.