The Edition
Narrative Report Intelligence
VOL. VIIBriefing01 / 07

Editor’s note · what you are looking at

An agentic reporting system that runs in production. This edition is its synthetic demo.

Most business reporting hands you a fixed grid of charts and asks you to hunt through it for what matters. This inverts that. A desk of specialist agents investigates the data, forms hypotheses, runs disproof tests, holds claims that are not yet earned, and publishes an edition — lead story first, evidence attached, reasoning open to challenge.

The real system runs against a national tire and auto service retailer’s data, on Snowflake. Everything you read below is invented — the company, the numbers, the analyst notes — so that the workflow can be shown without exposing client data.

How it works, and what is production versus roadmap →

VOL. VII · NO. 138 NARRATIVE REPORT INTELLIGENCE TUESDAY, JUNE 2, 2026

Operations · Lead Story

Three watch-list signals need operating proof before escalation

A real movement, or a measurement artifact? The investigator desk says probably both. The more interesting story is not that three signals appeared, but that each one stops just short of becoming an operating conclusion.

agents_sdk_primary gpt-5.5 thinking eval 100 quality 100 20 citations semantic retry passed

The five things · what to leave with

  1. Tire demand signals are present, but completed RO and utilization data are needed before calling a capacity squeeze .
  2. Pine Coast Paid Social softness is plausible, but attribution to the channel is not proven .
  3. Northeast Fleet may be influencing mix, but current evidence could reflect isolated or temporary work .
  4. Semantic retry downgraded overclaims and preserved the evidence trail.
  5. Learning memory now carries forward the proof standards for future Snowflake-backed runs.

Readable like a reporter. Auditable like a workpaper.

Every material statement is paired with an evidence path. Notes and point comments can add color and hypotheses, but they cannot carry numerical claims unless the source data agrees.

The lead

On the surface, the week looks like a manager's early-warning dashboard: tire pressure is rising, Paid Social looks soft in Pine Coast, and Northeast Fleet revenue is moving faster than bookings. But the agent desk found a more useful pattern underneath. Each signal is visible enough to deserve attention, and each is incomplete enough to punish a premature conclusion.

The result is a briefing that deliberately resists drama. The system is not saying the business is in trouble. It is saying the next operating conversation should be narrower, better evidenced, and easier to prove or disprove. That distinction matters: a weak report would convert movement into certainty; this one turns movement into a testable agenda.

Proof agenda

What the next query has to settle

01

Tires

Signal: Appointments and capacity index move together .

Missing: Completed ROs, utilization, cancellations, no-shows, and technician availability.

Decision use: Watch staffing and slot availability, but do not call a capacity squeeze yet.

02

Paid Social

Signal: Reach holds while appointments soften .

Missing: Spend, impressions, CTR, CPL, lead quality, and appointment conversion by source.

Decision use: Inspect funnel efficiency before changing budget or blaming creative fatigue.

03

Fleet

Signal: Revenue outruns appointment movement .

Missing: Customer concentration, repeat cadence, labor hours, ticket size, and job category mix.

Decision use: Treat as a mix question until repeat work proves durable demand.

Evidence points to live business questions, not firm conclusions.

The clean takeaway: management has three signals worth tracking, but the file does not yet clear the bar for escalation . The desk should keep them on watch until conversion, utilization, funnel, and repeat-work data close the gap.

That is a useful management posture. It keeps the agenda sharp without overstating the business case. The next meeting should not ask, “Are tires constrained?” It should ask whether completed work, service capacity, and missed demand agree with the appointment signal.

Demand signal, not a proven capacity constraint.

Repeated tire-related signals suggest momentum . But appointments and revenue can move for reasons that are not operational capacity: price, mix, scheduling behavior, or work that never converts into completed repair orders.

The operational read is therefore cautious. If completed ROs rise with appointment pressure and utilization tightens, this becomes a staffing and slot-management story. If completed work does not move, the signal may be scheduling friction, pricing, seasonal browsing, or demand that is not converting.

hold for completed ROscapacity index 88.6

Softness is plausible, but attribution remains unproven.

The file supports a possible Paid Social slowdown . It does not prove channel underperformance without spend, impressions, CTR, CPL, source-level conversion, and appointment trend separation.

The story to test is funnel quality, not vague marketing weakness. If spend and reach are stable while click-through and lead-to-appointment conversion fall, creative fatigue becomes plausible. If spend changed, source mapping drifted, or appointment availability tightened, the channel may be taking blame for an upstream or downstream issue.

creative fatigue hypothesisneeds funnel proof

Possible mix movement, not yet durable.

Fleet revenue may be influencing mix , but the current support could still be a one-off customer, isolated job, or ticket-size effect. Repeat cadence and customer concentration matter here.

The next query should split the movement by customer, job type, and ticket size. A broad fleet shift would show repeated customers and recurring categories. A fragile signal would concentrate in one account, one work order type, or one unusually large ticket.

watch customer concentrationrepeat work needed

The first draft wanted a cleaner story than the evidence allowed.

The semantic retry changed the report's center of gravity. The failed version leaned toward declarative claims: tires constrained, Paid Social underperforming, Fleet mix shifting. The accepted version keeps those ideas as hypotheses and names the tests that would make them publishable.

This is the agentic loop doing visible editorial work. It did not merely polish language; it changed claim strength, preserved citations, and left an audit trail for analysts who want to inspect the rejected path.

3 overclaims rejectedclaim strength reducedproof tests added

Point commentary

Comments become feedback for the next agent loop

Comments are saved as point-level context, then classified before they influence future runs. In production they would be written with section ID, paragraph ID, citation IDs, author, timestamp, feedback category, revision intent, and disposition.

Run desk

Watch the agentic loop execute

Idle0%
    1Signal Scoutscours metric rows, notes, checks
    2Hypothesis Generatorforms competing explanations
    3Investigation Plannerdefines disproof tests
    4Root Cause Analystchecks contradictions
    5Evidence Reviewerforces unsupported claims into retry
    6Narrative Editorwrites the story with caveats
    7Managing Editorapplies story budget and memory

    Evidence search

    The desk scours, narrows, and preserves the trail

    Search queue

      Evidence ledger

      Hypotheses under test

      Retry logic

      Where weak drafts are caught before they reach the reader

      Attempt 1

      Rejected overclaim

      “Tires are capacity constrained.”

      • Missing completed RO proof
      • No cancellation/no-show split
      • Revenue could be price or mix
      Semantic retry

      Evidence Reviewer sends it back

      The retry instruction requires the writer to separate signal detection from causal conclusion, preserve citations, and name the disproof tests.

      Rewrite as a watch-list signal. Keep citations. State what would validate or disprove the claim.
      Attempt 2

      Passed editorial gate

      “Tire demand signals are present, but completed RO and utilization data are needed before calling a capacity squeeze.”

      • Claim strength reduced
      • Evidence preserved
      • Next checks explicit

      Learning mechanism

      Run memory becomes habits of success

      The system records the causes of retries and successful editorial corrections. Next runs start with those habits loaded into the agent desk.

      Lessons carried forward

        Next-run behavior changes

        • Capacity narratives require completed ROs and utilization before escalation.
        • Paid media attribution requires spend, impressions, CTR, CPL, and conversion separation.
        • Fleet growth claims require repeat cadence and customer concentration checks.
        • Data-quality flags are promoted before any affected metric is used in a headline.

        Improvement scorecard

        Unsupported claims3 -> 0
        Named disproof tests2 -> 9
        Citation coverage100%
        Editorial quality100

        Trace history

        Loop history, queries, workpapers, and reasoning summaries

        The report can stay clean for readers, while analysts can open the production record behind it: what each agent searched, what it found, what it rejected, and what changed after retry.

        Run 01: evidence desk

        Retry packet

        Rejected draft claim

        Tires are capacity constrained; Paid Social is underperforming; Northeast Fleet has shifted mix.

        Reviewer reasoning summary

        The draft converted weak signals into causal conclusions. The reviewer required completed ROs for tires, source-level funnel metrics for Paid Social, and repeat-customer evidence for Fleet before escalation.

        Retry instruction
        Rewrite as evidence-bound watch-list signals. Preserve citations. Name what would validate or disprove each claim. Do not use appointment or revenue movement as causal proof.

        Learning memory written after run

        • Capacity claims must pass completed RO, conversion, cancellation/no-show, and utilization checks.
        • Marketing attribution must separate budget, reach, CTR, CPL, and conversion before blaming creative or channel quality.
        • Fleet movement must test concentration and repeat cadence before calling a durable mix shift.
        • Data-quality blockers are promoted before any affected metric can support the story.

        Expandable footnotes

        Citations with graphical support

        These are the evidence cards the story can lean on. Each figure is drawn with D3 and annotated for the exact point the narrative needs.

        [1] Display reporting caveat in Desert ValleyData quality

        Display has repeated partial or missing rows, so movement should remain directional and should not become a headline support point.

        synthetic_expansion:channel:Desert Valley:Display
        Display data quality flags

        0 = ok, 1 = partial, 2 = missing. The annotation highlights why the editor should avoid overclaiming Display revenue.

        [4] Tire capacity and appointment pressureService demand

        Tire appointments and capacity pressure moved together. The claim still needs completed repair orders before it can become a capacity conclusion.

        service_line:Tires:Lakeview:latest_week
        Tire appointments vs capacity index

        The supporting figure shows pressure, not proof. That distinction is the point of the story.

        [6] Pine Coast Paid Social funnel softnessMarketing funnel

        Impressions can look healthy while appointment yield weakens. The channel story needs source-level funnel metrics before attribution is fair.

        channel:Paid Social:Pine Coast:latest_week
        Paid Social reach vs appointments

        The divergence supports investigation, not a final diagnosis.

        [8] Northeast Fleet mix movementRevenue mix

        Fleet revenue moved faster than bookings, which may imply mix or average-ticket movement rather than broad demand growth.

        service_line:Fleet Service:Northeast:latest_week
        Fleet revenue vs appointments

        The desk should inspect RO count, customer concentration, labor hours, and repeat cadence next.

        Run archive

        Older runs, weekly drafts, and revision history live here

        In this demo, run records are saved in browser storage. In production, this page reads from the database: run ID, cadence, prompt, agent trace, citations, draft version, analyst comments, Ask This Report exchanges, retry packets, and publish status.

        Analyst notes

        Human context that can augment the run

        Notes are allowed to shape hypotheses and follow-up checks. The workflow treats them as context, not proof, unless matched to metric evidence.

        Deep dive observability

        Runtime, handoffs, tools, tokens, retries

        119,442msTotal runtime
        19,955Total tokens
        3,338Reasoning tokens
        1Semantic retry

        Tool calls

        load_company_profile, load_channel_rows, load_service_rows, load_analyst_notes, run_deterministic_checks, get_check_results, get_recent_analyst_notes, get_source_summary, compose_briefing

        Admin

        Schedule routine runs

        This hosted demo panel shows the intended operating control: a user can set cadence, instruction, and review mode for scheduled briefings.

        About

        This runs in production. What follows is the synthetic version of it.

        The reporting system on this site is a working demonstration built on invented data for a fictional company. The system it demonstrates is real: a recurring agentic briefing running in production against a national tire and auto service retailer’s performance data, on Snowflake. Everything under In production today below describes that live system. Everything under Designed, not yet built is roadmap, and is labelled that way deliberately — a page where every feature is production is a page you should not trust.

        In production today

        Snowflake is the evidence layer

        Every material number in a published report resolves back to Snowflake. Agents do not freely invent SQL against it. The warehouse is the source of truth for the claims, and the claims are bound to it rather than derived from a summary of it.

        The agent harness is the orchestration runtime

        Not an IDE assistant — the harness runs as the execution layer, and the agents inside it reach Snowflake through a governed binding maintained by our internal platform. The binding is what makes this safe to run on a schedule: agents get a controlled connection, not credentials and an open SQL prompt.

        Cortex Agents and Cortex Analyst are called as tools

        The harness calls Cortex Analyst against a semantic layer and Cortex Agents for the analytical work, treating their responses as tool output inside a longer loop it controls. Snowflake owns the question-to-answer step. The harness owns the investigation around it — which question is worth asking, whether the answer contradicts last period, and whether the evidence is strong enough to publish.

        Runs execute and log through the OpenAI Agents SDK

        Specialist agents, handoffs, tool calls, and traces are captured as first-class records inside an internal application, not as console output. The desk includes Signal Scout, Hypothesis Generator, Investigation Planner, Root Cause Analyst, Evidence Reviewer, Narrative Editor, and Managing Editor.

        The loop retries until it clears its bar

        Weak drafts are retried rather than published. Missing citations, unsupported numbers, data-quality blockers, and thin causal claims each send the loop back around. It publishes when the success criteria are met, and holding a claim is an acceptable outcome.

        The report and its trace publish together

        The result is published in an application alongside its logs. The trace is not debugging telemetry filed somewhere else — it ships with the story, because a claim a reader cannot interrogate is a claim they have to take on faith.

        What it has actually delivered

        Digestion, not discovery

        This has not been a discovery engine, and it would be dishonest to sell it as one. It surfaces findings a good analyst would have reached eventually by working through the dashboard. What changed is the digestion: the scanning, the ranking, the evidence assembly, and the defensible writing-up — the work that consumed analyst hours and happened live in meetings — now happens before anyone opens anything.

        That is information management, not clairvoyance. A system that reliably produced startling revelations every period would not be investigating; it would be performing, manufacturing a story to justify the run. The bar that makes this trustworthy is the same bar that keeps it unspectacular.

        Designed, not yet built

        Editable, versioned drafts

        Human edits should create new versions, preserve the generated draft, and trigger citation-integrity checks when a user introduces new claims or numbers.

        Point comments as context

        Comments attached to a specific claim, paragraph, chart, or citation, stored with their target and evidence references, so revision agents can retrieve them as context rather than as unbounded chat.

        Feedback as a classified loop

        Comments are signals, not instructions. The loop should classify feedback, decide whether it becomes a prompt rule, eval case, evidence request, preference, backlog item, or no action, then expose that disposition in the run history.

        Clickable SQL evidence

        Superscript markers on material claims opening an evidence drawer: query spec, compiled SQL, source rows, freshness status, chart spec, and the trace events from the agent that used it.

        Learning memory across runs

        Recording habits of success — proof standards, source reliability, editorial preferences, failed claim patterns — and retrieving them before the next run starts.

        Ask This Report

        A marginal, grounded conversational layer: highlight a passage, ask a follow-up, get an answer constrained to report context, citations, evidence, notes, and run history — without letting chat replace the crafted narrative.

        Full operational layer

        Authentication, role-based access, approval workflows, trace redaction, query cost limits, and eval harnesses, so analysts can dive into the machinery beneath an edition-quality report.

        Page experience

        Edition pages rather than a single scroll. Transitions like paper being laid down: folios update, rules and headlines enter first, evidence unfolds from the margin, the agent timeline moves like a production desk.

        What we would want from the platform next

        Cortex Analyst answers; it does not remember

        Analyst is very good at turning a question into a defensible answer against a semantic view. What it does not do is carry state across invocations — notice that this period’s answer contradicts last period’s, or that a claim it just made was held for thin evidence. We built that loop ourselves. A first-class notion of run continuity would remove the most-copied piece of glue code in systems like this.

        Semantic view synonyms are the highest-leverage surface, and the least documented

        Users say “non-brand”; the column says IS_BRANDED. Every gap like that is a question Analyst answers worse. Synonyms and comments are not documentation — they are the input the model reasons over, and tuning them beat every prompt change we tried. The docs treat them as a nice-to-have. A worked example with deliberately bad synonyms beside good ones, same question asked of both, would make the case in thirty seconds.

        Trace continuity across the boundary

        Our orchestration traces live in one system and Cortex’s execution detail lives in another. Stitching a published claim back through the loop and into the query that produced it is work we do by hand. A correlation ID that survives the call into Cortex Agents would make end-to-end provenance close to free — and provenance is the whole product here.

        Model names are string literals with no safety net

        A wrong model name is a runtime error you hit after writing the query, and the roster moves, so older examples fail on paste. An error that named the available models in the region would eliminate the entire category. The information is clearly present at the point of failure.

        Say the grain thing out loud

        Classify at the distinct grain, not the row grain. On a real export that is a 20× cost difference, and it is also a correctness issue — at row grain, identical text can receive different labels on different rows and nothing errors. It is a decision made in the first ten minutes and painful to retrofit. It deserves to be in the first tutorial, not learned on a bill.

        A reference loop with the guardrails already in it

        The gap between “here is how to call a model” and “here is how to safely run what it wrote” is where teams either stall or, worse, do not stall and ship something unbounded. A canonical bounded loop — read-only role, allowlist, row caps, retry limits, stop conditions — would raise the floor across the ecosystem.

        Boundary. This public version contains no client data, production credentials, proprietary source systems, private prompts, or live operational workflows. The company, the data, and the analyst notes are invented. The architecture, the workflow, and the editorial discipline are the real ones.