WingNews logo WingNews
top | new | best | ask | show | jobs
top | item 43595930

(no title)

nattaylor | 11 months ago

Is pre-training in FP8 new?

Also, 10M input token context is insane!

EDIT: https://huggingface.co/meta-llama/Llama-3.1-405B is BF16 so yes, it seems training in FP8 is new.

discuss

order

jumpCastle|11 months ago

Deepseek v3 was FP8
powered by hn/api // news.ycombinator.com