Trace capture
Collect inputs, model outputs, retrieved context, tool calls, errors, latency, and cost for each representative workflow.
LLM evaluation platform
A production LLM evaluation platform should compare model, prompt, retrieval, and tool changes against the workflows users actually run. Vectory focuses on traces, evidence, stop conditions, and release decisions.
An LLM evaluation platform should score final answers and the trace that produced them. The trace shows whether the system retrieved useful evidence, used tools with discipline, recovered after failures, and verified completion.
Collect inputs, model outputs, retrieved context, tool calls, errors, latency, and cost for each representative workflow.
Enforce required observations, source use, tool budgets, approval gates, stop conditions, and completion verification.
Route high-judgment failures to calibrated human or LLM-as-judge review with stable rubrics.