ROCKMAN / v0.2

Rockman

An evaluation benchmark for software engineering and reasoning tasks, designed around controlled splits, repeatable evaluation, and protected test material.

INSERLOFT / BENCHMARK
01 / BENCHMARK STATE
220Total tasks
154Public
44Private
22Hidden
02 / TASK FIELD
A benchmark is not a list of questions. It is a measurement environment.
DIFFICULTY
1 → 7

Tasks are organized across difficulty levels and categories. Results are aggregated only after individual evaluations have completed.

The public split is for development. Protected splits are for evaluation integrity.

03 / PROTOCOL
From generation to statistical evidence.
INPUTTask prompt
RUNModel output
CHECKEvaluation
AGGREGATEPass@k
REPORTConfidence + significance
04 / INTEGRITY
Public enough to develop. Protected enough to evaluate.
PROTECTED MATERIAL
AES-256-GCM

Rockman v0.2 separates public, private, and hidden evaluation material. Hidden tests are encrypted and integrity checked to reduce accidental exposure and contamination.

SHA256 / 3352bcf…

05 / RESULTS
Results should be reproducible, inspectable, and comparable.
REPORTS
CSV + PNG

Pass@k · Wilson confidence intervals · statistical significance · difficulty analysis · category analysis

Use the official benchmark configuration when comparing models. Modified protocols should be identified as custom evaluations.

06 / REPOSITORY
Read the benchmark. Run it. Inspect the protocol.
GITHUB ↗   HUGGING FACE ↗