top | item 46337386

(no title)

cg5280 | 2 months ago

I like the inclusion of the graph at the end to compare progress. It would be cool to compare this directly to competing models (Claude, GPT, etc).

discuss

kqr|2 months ago

It would unfortunately also need several runs of each to be reliable. There's nothing in TFA to indicate the results shown aren't to a large degree affected by random chance!

(I do think from personal benchmarks that Gemini 3 is better for the reasons stated by the author, but a single run from each is not strong evidence.)

casey2|2 months ago

TFA says multiple times that the results are affect by random chance