WingNews

dwohnitmok|11 days ago

Not anymore. This benchmark is for LLM chess ability: https://github.com/lightnesscaster/Chess-LLM-Benchmark?tab=r.... LLMs are graded according to FIDE rules so e.g. two illegal moves in a game leads to an immediate loss.

This benchmark doesn't have the latest models from the last two months, but Gemini 3 (with no tools) is already at 1750 - 1800 FIDE, which is approximately probably around 1900 - 2000 USCF (about USCF expert level). This is enough to beat almost everyone at your local chess club.

runarberg|11 days ago

Wait, I may be missing something here. These benchmarks are gathered by having models play each other, and the second illegal move forfeits the game. This seems like a flawed method as the models who are more prone to illegal moves are going to bump the ratings of the models who are less likely.

Additionally, how do we know the model isn’t benchmaxxed to eliminate illegal moves.

For example, here is the list of games by Gemini-3-pro-preview. In 44 games it preformed 3 illegal moves (if I counted correctly) but won 5 because opponent forfeits due to illegal moves.

https://chessbenchllm.onrender.com/games?page=5&model=gemini...

I suspect the ratings here may be significantly inflated due to a flaw in the methodology.

EDIT: I want to suggest a better methodology here (I am not gonna do it; I really really really don’t care about this technology). Have the LLMs play rated engines and rated humans, the first illegal move forfeits the game (same rules apply to humans).

cesarvarela|11 days ago

Yeah, but 1800 FIDE players don't make illegal moves, and Gemini does.

deadbabe|11 days ago

Why do we care about this? Chess AI have long been solved problems and LLMs are just an overly brute forced approach. They will never become very efficient chess players.

The correct solution is to have a conventional chess AI as a tool and use the LLM as a front end for humanized output. A software engineer who proposes just doing it all via raw LLM should be fired.

overgard|11 days ago

They have literally every chess game in existence to train on, and they can't do better than 1800?

nicowesterdale|4 days ago

I wrote a, I hope, amusing breakdown of the structural reasons why off-the-shelf Large Language Models physically cannot "see" a chess board, and continue to make illegal moves, and teleport pieces as seen in Gotham Chess' latest videos.

https://www.nicowesterdale.com/blog/why-llms-cant-play-chess

iugtmkbdfil834|11 days ago

Hm.. but do they need it.. at this point, we do have custom tools that beat humans. In a sense, all LLM need is a way to connect to that tool ( and the same is true is for counting and many other aspects ).

Windchaser|11 days ago

Yeah, but you know that manually telling the LLM to operate other custom tools is not going to be a long-term solution. And if an LLM could design, create, and operate a separate model, and then return/translate its results to you, that would be huge, but it also seems far away.

But I'm ignorant here. Can anyone with a better background of SOTA ML tell me if this is being pursued, and if so, how far away it is? (And if not, what are the arguments against it, or what other approaches might deliver similar capacities?)

menaerus|11 days ago

Did you already forget about the AlphaZero?

BeetleB|11 days ago

Are you saying an LLM can't produce a chess engine that will easily beat you?

emp17344|11 days ago

Plagiarizing Stockfish doesn’t make me good at chess. Same principle applies.

(no title)

discuss