Kyalulu/docs
GitHub JA
Browse guides⌄

Build & evaluate

What CharacterBench measures

Use a fixed 12-character, 264-response pilot while separating mechanical checks from character quality.

At a glance

  • CharacterBench is an independent MIT-licensed pilot for evaluating Japanese Character AI.
  • Twelve characters, 120 single-turn probes and twelve 12-turn dialogues produce 264 responses per model per repeat.
  • Structured-state and surface-constraint scores are not overall character-quality scores or a validated leaderboard.
On this page
  1. At a glance
  2. Separate from the Runtime benchmark
  3. Inspect the setup without inference
  4. Separate mechanics from character quality
  5. Before comparing real models

Separate from the Runtime benchmark#

CharacterBench v0.1 is an independent MIT-licensed engineering pilot for Japanese character AI. Twelve original adult fictional characters, 120 single-turn probes, and twelve 12-turn dialogues produce 264 responses per model per repeat. It is a fixed public synthetic corpus, not a human-validated model leaderboard.

benchmarks/official is a separate track using Kyalulu Runtime state, memory, persistence, and shared prompts. CharacterBench Core and System tracks must also remain separate. The bundled System bridge is a reference session adapter, not integration with the actual Kyalulu app. MIT applies to this subdirectory's code and newly authored corpus; it does not change the main repository's AGPL license.

Inspect the setup without inference#

Use Python 3.10+; runtime dependencies are standard-library only. Work from benchmarks/characterbench. Data validation and call planning need neither a model nor an API key.

Set-Location benchmarks/characterbench
python -m kcb validate
python -m kcb plan --suite all --repeats 3
# Optional fixed-MOCK report; use a new output directory
python -m kcb demo --out runs/my-demo

plan makes no network calls. One all repeat produces 264 responses, three produce 792, and smoke produces 22. Free-text evaluation has 84 units for all and seven for smoke; a 12-turn dialogue is one evaluation unit. demo checks the workflow with fixed MOCK responses, not real-model speed or quality. The repository version does not include the ZIP distribution's HTML launchers.

Separate mechanics from character quality#

The main mechanical metrics from run are strict-JSON state diagnostic accuracy and per-field errors. Literal constraints such as length or bullet formatting can also be checked, but keyword matches do not establish personality appeal, natural Japanese, or story quality. Semantic evaluation requires a separate LLM rubric judge, order-reversed A/B comparisons, or blind human review.

LLM judges are uncalibrated. Verifying that a quoted passage exists does not prove the judgment is correct. Read evaluated counts, missing judgments, judge errors, and order sensitivity together. A small group's preferences or a Gemma pilot on one machine cannot establish an overall ranking. Match Core/System tracks, data hashes, protocols, and repeat counts, and retain generation and judging conditions separately.

Before comparing real models#

  1. Start an OpenAI-compatible chat server and load its model yourself; the tool does not download or start them. Use doctor to check the exact model ID and begin with smoke. --probe and run involve real generation.
  2. Record weights, quantization, backend version, thinking, MTP, and sampling. Use a new output directory when conditions change.
  3. Retain responses, errors, missing measurements, and the manifest. Completion alone does not establish quality. Token-truncated or abnormally finished responses remain available for audit but are not scored or used in subsequent dialogue.
  4. Use --resume for the same conditions and explicitly add --retry-errors only to retry failures. Successful responses are not regenerated to cherry-pick better results.

Without usage, token counts and costs are unmeasured, not zero. No observable reasoning output does not prove the absence of internal processing. regrade evaluates stored responses without new inference and records generation and grading versions separately. Review raw inputs and history before publication, and never give human raters the model identity key. See the development guide for Runtime responsibilities or the documentation index for all guides.

Sources for this article

Edited from public GitHub materials. Links are pinned to the reviewed commit.

Search the docs

↑ ↓ select · Enter open · Esc close