Nous Research is an open-source AI research lab known for its Hermes models, and also runs Psyche, a decentralized AI-training network coordinated on Solana, backed by a roughly $50M Series A led by Paradigm. This role focuses on agent capability evaluations, benchmark design, LLM-as-judge systems, failure analysis, and evaluation infrastructure.
The engineer will run the full evaluation pipeline end-to-end and own the recurring evaluation workflow that informs model development.
A remote evaluation-focused ML role at a well-funded open-source AI lab that also runs a decentralized, Solana-coordinated training network.

Nous Research is an open-source AI research lab known for its Hermes models, also running Psyche, a decentralized AI-training network coordinated on Solana. Backed by a ~$50M Series A led by Paradigm at a near-$1B valuation.
Apply To This Job<<>>
Support us by letting the company know you found them on our website.
Go To the Offer