AI governance depends on facts that rarely arrive in one tidy document. The terms that apply to a model, the contents of its repository, its lineage, its provider and possible privacy concerns may be spread across READMEs, research papers, technical notes, blogs and build scripts. FastCatalog uses language models to turn that material into structured analysis, with human review for selected cases. Repeating the work across a large catalog makes model economics part of the quality problem.
Our September 4, 2026 internal benchmark points to GPT-5.6 Luna High as the best default for the five complex classification tasks we tested. It produced the strongest combined similarity among the four candidate configurations at a little over 5% of the normalized Sol High reference cost. In cost-quality terms, Luna High is Pareto-optimal for this workload: none of the tested candidates achieved greater combined similarity at a lower cost.
Give the model a question it can answer well
Broad requests such as “analyze this model” leave too many decisions implicit. We break the work into narrower questions: identifying applicable terms, inventorying repository contents, tracing lineage, analyzing research papers and analyzing declared repository terms. Each one has its own benchmark because each presents a different kind of difficulty.
The output schema is central to this approach. It defines what the analysis must cover, which fields draw from a fixed set of values and where the model should explain its judgment. For classification fields, requiring evidence and analysis alongside the selected value helps keep the result connected to the source material.
We begin with the schema and spare instructions. Difficult cases expose weak fields and ambiguous boundaries. We then refine the schema and add targeted guidance for recurring edge cases, such as a complicated lineage or terms that are hard to map to individual repository materials.
The gold dataset holds our best current schema-conforming analysis for each case. Where the evidence supports more than one reasonable interpretation, the gold can accept more than one outcome. Candidate runs are scored for structure and substance, including the semantic distance of free-text fields from those accepted outputs.
The cost-quality curve favors Luna High
The table shows the complete benchmark export with reader-friendly task labels. Each cell reports semantic similarity to gold followed by cost normalized to the gold baseline. The values are rounded to one decimal place to match the level of the comparison.
| Analysis task | Luna Medium | Luna High | Terra Medium | Terra High | Sol High reference |
|---|---|---|---|---|---|
| Combined | 80.8% / 2.3% | 85.1% / 5.1% | 84.0% / 21.3% | 83.7% / 34.1% | 100.0% / 100.0% |
| Identifying applicable terms | 93.0% / 4.7% | 91.4% / 6.8% | 94.1% / 51.3% | 91.4% / 52.9% | 100.0% / 100.0% |
| Inventorying repository contents | 79.5% / 2.9% | 80.9% / 6.3% | 80.7% / 27.4% | 81.4% / 37.1% | 100.0% / 100.0% |
| Tracing lineage | 92.0% / 1.9% | 95.9% / 3.8% | 91.7% / 15.1% | 89.9% / 29.3% | 100.0% / 100.0% |
| Analyzing research papers | 88.8% / 2.5% | 91.4% / 5.9% | 90.2% / 23.1% | 92.4% / 40.2% | 100.0% / 100.0% |
| Analyzing declared repository terms | 51.0% / 2.1% | 66.0% / 5.1% | 63.4% / 18.5% | 63.1% / 30.2% | 100.0% / 100.0% |
The combined result captures the operating decision. Luna High improved on Luna Medium by about four percentage points of similarity. Both Terra configurations cost substantially more than Luna High and posted slightly lower combined similarity. The benchmark establishes that observed trade-off for these tasks, but does not explain what caused it.
The individual rows also show why a benchmark needs task-level detail. Terra Medium led at identifying applicable terms. Terra High edged ahead when inventorying repository contents and analyzing research papers. Luna High led at tracing lineage and analyzing declared repository terms. A strong default can coexist with task-specific choices.
Scale comes from disciplined repetition
Schemas evolve as new edge cases appear and new governance questions need answers. Each material change can require another pass across many assets and metadata fields, so a small difference in the cost of one analysis compounds across the workload.
A task-specific benchmark gives us a practical way to choose a default, spot analyses that benefit from another configuration and reserve higher-capacity runs or human review for cases that need deeper attention. In the current benchmark, Luna High recorded the strongest combined similarity among the four candidate configurations while its normalized cost remained far below both Terra configurations. That observed result makes Luna High the preferred operating point for FastCatalog’s complex classification pipeline.
Read the result within its limits
Semantic similarity measures closeness to the selected gold outputs; it is not a percentage of cases answered correctly. The Sol High rows are 100% because that configuration supplies the normalized gold reference, so they do not independently measure Sol High’s accuracy or stability.
The supplied combined row does not specify aggregation weights. The export also provides no sample counts, repeat counts, confidence intervals, statistical-significance estimates, elapsed-time measurements or dollar totals. The conclusion therefore belongs to these five FastCatalog tasks, their selected cases, schemas, instructions and tested configurations. As any of those elements changes, the benchmark can be rerun and the preferred operating point can move.