AI EngineeringEvaluating AI Agents in CI: Scoring Full Trajectories, Not Just the Final Answer
Final-answer assertions keep your CI green while your agent calls the wrong tool, loops extra steps, and quietly burns budget. Here is the trajectory-scoring eval harness I run: deterministic tool mocks, six scored dimensions, and an LLM judge kept on a rubric and audit.










