Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Combined with our latest 30T-token multimodal pre-training corpus [...]

Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: