Stripe released a new benchmark suite to test how effectively AI agents perform end-to-end software engineering tasks. While models show promise in code generation, the research highlights significant weaknesses in state management, verification, and error recovery during complex production scenarios.