We are back with our interview series and a very special guest today! Anastasios Angelopoulos has been at the center of how the industry measures model quality in the wild. We discuss the origins of Chatbot Arena, what an Arena score actually captures, how preference and factuality should (and shouldn't) be combined, Agent Arena's performance'cost frontier, AutoEval, and what evaluation looks like once harnesses and tools enter the ranking. I'm Anastasios, co-founder and CEO of Arena. My path in AI has been pretty nonlinear. I started working in AI + medicine, then came to believe that one of the most important problems deploying AI in the medical sector is reliability. That led me to do a PhD at Berkeley focused on the theoretical foundations of AI reliability. In the process, I worked on many projects, one of which was Chatbot Arena, which evolved into Arena. Now, we are building an AI evaluation company that measures the performance and reliability of AI in real-world workflows...
learn more