Results for "SWE benchmarks"
Keyword scan across titles, descriptions, summaries, and tags. For interview listings, try Guest appearances.
1 result
Episodes
StandardSummaries only[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Latent Space: The AI Engineer Podcast· Dec 31, 2025
From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard fo…
openaianthropicevalsmultimodal