I have the feeling that Kiro is more efficient. Is it just me, or is it real?
I had the feeling that Kiro lasted longer than Codex or Claude Code, but without numbers not even I believed it. So I built a benchmark with the same models, the same exercises and the same prompt across all three tools. The results surprised me, and not in the direction I expected.