πŸ“Š

LLM Eval Framework

Admin

βš™οΈ Configuration

Loading…
When off, no LLM judge calls are made and scores stay null.
100%
% of queries that get evaluated. 100% = every query.
0.70
Queries scoring below this are marked flagged in the table.
Model used for judge calls. Should be a fast, cheap model.
πŸ¦™
Ollama
☁️
AWS Bedrock
πŸ€–
Claude API
πŸ”·
Azure OpenAI
βœ“ Saved

πŸ“ˆ Summary

β€”
Total Queries
β€”
Evaluated
β€”
Avg Overall
β€”
Avg Relevance
β€”
Avg Faithfulness
β€”
Avg Context Prec.
β€”
Flagged
β€”
Hallucinations

πŸ—‚ Eval Scores

Question Overall Relevance Faithfulness Context Halluc. Status Time
Loading…
Page 1

πŸ—ƒ Golden Dataset

Question Expected Answer (excerpt) Retrieval Actions
Loading…
+ Add Q&A Pair
βœ“ Added

πŸƒ Eval Run History

Started Triggered by Status Model Questions Avg Score vs Baseline Flagged
No eval runs yet. Click "Run Eval Now" to start.
Page 1