The prompt
Build a scoring rubric for evaluating LLM outputs on the task below.
Dimensions: accuracy, completeness, instruction-following, safety, and one task-specific dimension.
For each dimension: a 1-5 scale with concrete anchors for 1, 3, and 5.
Then state whether to use an LLM-as-judge or human judge for each dimension, and why.
Task: {{task}}Replace the {{fields}} with your own context and tighten the rules to match your domain.


