quote: Proving code works with math instead of tests used to cost too much at scale.
NEAR AI is now #1 on Lean Eval v1, | Hanami
quote: Proving code works with math instead of tests used to cost too much at scale.
NEAR AI is now #1 on Lean Eval v1, Lean's own benchmark of extremely hard formalization problems.
It runs almost entirely on open-weight DeepSeek V4.1 Flash, so it's cheap.
Next: NEAR contracts. https://x.com/ilblackdragon/status/2106094405564408006 | NEAR AI has claimed first place on the Lean Eval v1 leaderboard.
Lean Eval v1 is an active benchmark maintained by Lean FRO that contains extremely hard formalization problems. We reached first place despite joining the competition just four weeks ago, a full month later than other top participants.
NEAR is very heavily investing in formal verification efforts. Our ultimate goal is to formally prove the protocol and all the core contracts, and to provide a practical tool for others to formally verify their contracts in a fast and affordable way.
The system that enabled us to claim first place is almost entirely powered by DeepSeek V4.1 Flash, and thus is (a) extremely cheap, and (b) has no dependencies on third parties because the model is open weight.
When it comes to formally verifying software, the approach boils down to three pillars:
1. Representing the program in a way that can be reasoned about;
2. Expressing the properties that we desire to prove;
3. Coming up with a proof and formalizing it in Lean.
Until recently, step (3) was prohibitively expensive for practical use at scale. With the system that we used for Lean Eval, we now have it solved. Stay tuned for updates on how we apply it in practice to formally verify NEAR contracts.