top | item 46786002

(no title)

kevmo314 | 1 month ago

Reading their paper, it wasn't trained from scratch, it's a fine tune of a Qwen3-32B model. I think this approach is correct, but it does mean that only a subset of the training data is really open.

discuss

No comments yet.