<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on Subhajai Benchadhikul</title><link>https://subhajai.benchadhikul.com/tags/evaluation/</link><description>Recent content in Evaluation on Subhajai Benchadhikul</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 09 Sep 2026 07:00:00 +0700</lastBuildDate><atom:link href="https://subhajai.benchadhikul.com/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>The Eval Gap: Nobody Can Tell You If It's Working</title><link>https://subhajai.benchadhikul.com/posts/2026-09-09-the-eval-gap-ai-features/</link><pubDate>Wed, 09 Sep 2026 07:00:00 +0700</pubDate><guid>https://subhajai.benchadhikul.com/posts/2026-09-09-the-eval-gap-ai-features/</guid><description>Most teams shipping AI features cannot answer the simplest question about them — is this actually good? Evaluation is not a testing chore. It is the definition of correct, and writing it down is the hardest and most valuable work in the project.</description></item></channel></rss>