Results for "SWE benchmarks"

Keyword scan across titles, descriptions, summaries, and tags. For interview listings, try Guest appearances.

1 result

Episodes

  • Latent Space: The AI Engineer Podcast
    StandardSummaries only

    [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

    Latent Space: The AI Engineer Podcast· Dec 31, 2025

    From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard fo

    openaianthropicevalsmultimodal