Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests. - The New Stack
Sierra has open-sourced Hyper-𝜏-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents.