← How it works

AI-generated interactive documents: components, validation, and continuous testing

Learn how OWL Compose helps AI turn content into documents, and how we check the results.

Approach comparison

Where OWL Compose has the advantage over direct AI generation and templates

Direct generation is fast, but the result depends on the model. Templates are steadier, but limit what can be made. OWL Compose gives the agent room to compose while keeping the finished work stable, editable, and publishable.

  1. Direct AI generation

    You get a result quickly, but each run still depends on how well the model performs.

    Fast resultHigh freedom
  2. AI templates

    Easy to start and more predictable, but content and layout stay inside the template’s limits.

    Easy startStable output
  3. OWL Compose

    Compose layouts, charts, and visuals freely, then keep editing and publishing the same work.

    Finished qualityKeeps evolving
Conceptual comparison

Six practical capabilities, with different strengths

Direct generation leads on speed and freedom; templates lead on predictability; OWL Compose is stronger across the whole workflow.

Direct AI generationAI templatesOWL Compose

Continuous testing

Good work should not require the most expensive model

We repeatedly test OWL Compose on baseline models near the price–performance frontier. The tasks and configuration stay fixed; when models improve, we rerun the tests and adjust the system so that capability reaches the finished work.

Adaptive abstraction

Give the model structure where it needs help — and freedom where it does not

The right authoring layer is not uniformly high-level or low-level. We set the boundary according to what baseline models can reliably create.

Hard for AI
Package the complexity

Advanced charts, spatial layouts, and coordinated interactions become tested components with a small set of meaningful parameters.

A few parameters → advanced, stable output
Easy for AI
Expose the atoms

When models can compose reliably, we split the capability into finer primitives instead of locking it inside a template.

Fine-grained syntax → more creative freedom
Recurring ablationRemove, merge, split, retest.

Using the same baseline models and fixed tasks, we regularly ablate each layer. A boundary stays only when the evidence shows it helps.

Price–performance frontier

15 model configurations · AA Index v4.2 · USD per index task

OpenAIZ AIDeepSeekMiniMaxGoogleMetaKimiAnthropic
GPT-5.6 Luna (max): $0.1, 43GLM-5.3-Flash: $0.18, 46DeepSeek V4 Flash 0731 (max): $0.14, 41MiniMax-M3: $0.23, 36DeepSeek V4 Pro 0813 (max): $0.33, 42Gemini 3.1 Pro Preview: $0.34, 37Gemini 3.7 Flash (high): $0.55, 45GPT-5.6 Terra (max): $0.81, 47Muse Spark 1.3 (max): $0.96, 53GPT-5.6 Sol (max): $1.25, 51GLM-5.3 (max): $1.26, 49Kimi K3 (max): $1.58, 50GPT-6 Astra (max): $2.57, 55Claude Opus 5 (max): $4.21, 54Claude Fable 5.1 (max, fallback): $6.12, 57GPT-5.6 Luna (max)GLM-5.3-FlashMuse Spark 1.3 (max)Claude Fable 5.1 (max, fallback)$0.1$1$1030405060Cost per task (USD · log scale)
Artificial Analysis measures general model intelligence, not a measure of OWL Compose work quality.

Colored points trace the upper envelope on a log-cost axis. Grey points cost more for the same or lower score; hollow points sit below the envelope. This is a sourced sample, not the full leaderboard. A frontier point is a candidate for our own work-quality tests, not proof of passing them. Missing task costs are excluded.

15 model configurations · AA Index v4.2 · USD per index task
ModelUSDAA v4.2Price–performance frontier
GPT-5.6 Luna (max)0.1043Baseline candidate
GLM-5.3-Flash0.1846Baseline candidate
DeepSeek V4 Flash 0731 (max)0.1441Dominated
MiniMax-M30.2336Dominated
DeepSeek V4 Pro 0813 (max)0.3342Dominated
Gemini 3.1 Pro Preview0.3437Dominated
Gemini 3.7 Flash (high)0.5545Dominated
GPT-5.6 Terra (max)0.8147Below the envelope
Muse Spark 1.3 (max)0.9653Baseline candidate
GPT-5.6 Sol (max)1.2551Dominated
GLM-5.3 (max)1.2649Dominated
Kimi K3 (max)1.5850Dominated
GPT-6 Astra (max)2.5755Below the envelope
Claude Opus 5 (max)4.2154Dominated
Claude Fable 5.1 (max, fallback)6.1257Baseline candidate