Model Evaluation Dashboard
Baymax Chatbot
Responses by the LLM are evaluated against a golden dataset using an LLM-as-a-judge, scored on accuracy, relevance, and conciseness.
How LLM responses are generated →Setup
Judge LLM
Groq (llama-3.1-8b-instant)
Golden dataset
12 questions
Eval Scores
Each answer is scored 1 to 5 using an LLM-as-a-judge.
Eval History
Loading...