Fleebs-Logo
Details werden geladen...

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases | Hacker News

Ähnliche Seiten

https://news.ycombinator.com/item?id=49124336

13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS | Hacker News

https://news.ycombinator.com/item?id=49124336
https://news.ycombinator.com/item?id=49076391

Benchmarking Opus 5 on SlopCodeBench | Hacker News

https://news.ycombinator.com/item?id=49076391
https://news.ycombinator.com/item?id=48839355

MIRA: Multiplayer Interactive World Models Trained on Rocket League | Hacker News

https://news.ycombinator.com/item?id=48839355
https://news.ycombinator.com/item?id=49601338

AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200 | Hacker News

https://news.ycombinator.com/item?id=49601338
https://news.ycombinator.com/item?id=48296359

Training our own AI models | Hacker News

https://news.ycombinator.com/item?id=48296359
https://news.ycombinator.com/item?id=49646778

Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1 | Hacker News

https://news.ycombinator.com/item?id=49646778