Model Evaluation Dashboard

Baymax Chatbot

Responses by the LLM are evaluated against a golden dataset using an LLM-as-a-judge, scored on accuracy, relevance, and conciseness.

How LLM responses are generated →

Setup

Judge LLM Groq (llama-3.1-8b-instant)
Golden dataset 12 questions

Eval Scores

Each answer is scored 1 to 5 using an LLM-as-a-judge.

Eval History

Loading...