WingNews logo WingNews
top | new | best | ask | show | jobs
top | item 42475724

(no title)

dimitry12 | 1 year ago

"Solver" is `meta-llama/Llama-3.2-1B-Instruct` (1B model, and they use 3B for another experiment), and verifier is `RLHFlow/Llama3.1-8B-PRM-Deepseek-Data`.

See https://github.com/huggingface/search-and-learn/blob/b3375f8... and https://github.com/huggingface/search-and-learn/blob/b3375f8...

In the original paper, they use PaLM 2-S* as "solver" and its fine-tune as "verifier".

discuss

order

No comments yet.

powered by hn/api // news.ycombinator.com