26 terms
Definitions in this topic
- ProductionAgent simulationAgent simulation explained: controlled environments for testing behavior before production.
- ProductionAgent trajectoryAgent trajectory explained: the ordered record of an AI agent's messages and tool calls, how to score it, and what we check in our own agent runs.
- Agentic AIAI agent evalsAI agent evals explained: end to end, trajectory, and component level testing, and why single-turn accuracy is not enough.
- ProductionAI evaluation harnessAI evaluation harness explained: the dataset, scorers, and pipeline that turn one off tests into a repeatable gate.
- BusinessAI ROI measurementAI ROI measurement done honestly: per workflow baselines, fully loaded costs, and why hours saved do not equal money saved.
- ProductionArize PhoenixArize Phoenix explained: OpenTelemetry-native LLM tracing and evals, self-hosting, and where it sits against Langfuse and LangSmith.
- ProductionBenchmark contaminationBenchmark contamination explained: test data leaking into model training, inflated scores, detection limits, and safer evaluation design.
- BusinessCost per taskCost per task explained: how to compute the real cost of agent work and use it in build, buy and pricing decisions.
- ProductionEval datasetEval datasets explained: representative test cases, expected outcomes, failure coverage, versioning, and separation from training data.
- ProductionJudge biasJudge bias explained: position, verbosity, style, and self-preference effects in LLM evaluation, plus practical detection tests.
- ProductionLLM as a judgeWhat LLM as a judge means, the biases judges carry, and how teams validate automated evaluation against human labels before trusting it.
- ProductionModel calibrationModel calibration explained: aligning predicted confidence with observed frequency, measuring reliability, and setting decision thresholds.
- ProductionModel validationModel validation explained: testing models on unseen data, and what the discipline demands in regulated industries.
- ProductionOnline vs offline evaluationOnline vs offline AI evaluation: when to test fixed datasets before release and when to monitor live production runs.
- ProductionPairwise evaluationPairwise evaluation explained: compare two AI outputs, define preference rubrics, handle ties, and control order and judge bias.
- LLM foundationsProcess supervisionProcess supervision explained: evaluating intermediate reasoning steps, how it differs from outcome supervision, and its labeling tradeoffs.
- ProductionPrompt Regression TestingPrompt regression testing explained: compare a changed LLM system against representative cases before it reaches production.
- ProductionRAG evaluationRAG evaluation explained: measuring retrieval relevance, answer correctness, faithfulness, and end-to-end usefulness.
- Agentic AIReflection Agent PatternReflection agent pattern explained: the draft, critique, revise loop, and when self-review improves an AI workflow.
- LLM foundationsReward modelReward models explained: learned preference scores used for optimization, how they are trained, and why reward hacking remains a risk.
- ProductionSimulated-user evaluationSimulated-user evaluation explained: automated conversational testing, scenario coverage, simulator bias, and validation against real users.
- ProductionSynthetic dataSynthetic data explained: algorithmically generated examples for training and testing, useful cases, privacy limits, and validation needs.
- ProductionTask Completion RateTask completion rate explained: an outcome metric for AI agents, and why success percentage alone is not enough.
- ProductionTrace-Based EvaluationTrace-based evaluation explained: evaluate an AI system's full execution path, not only the final answer.
- ProductionUncertainty estimationUncertainty estimation explained: data and model uncertainty, confidence signals, validation, abstention, and human escalation.
- ProductionVoice agent evaluationVoice agent evaluation explained: measure task success, speech quality, timing, interruptions, and recovery across realistic calls.