Weekly creates a distinct run with its own cadence, prompt, trace, citations, and draft version.
Editor’s note · what you are looking at
An agentic reporting system that runs in production. This edition is its synthetic demo.
Most business reporting hands you a fixed grid of charts and asks you to hunt through it for what matters. This inverts that. A desk of specialist agents investigates the data, forms hypotheses, runs disproof tests, holds claims that are not yet earned, and publishes an edition — lead story first, evidence attached, reasoning open to challenge.
The real system runs against a national tire and auto service retailer’s data, on Snowflake. Everything you read below is invented — the company, the numbers, the analyst notes — so that the workflow can be shown without exposing client data.
Operations · Lead Story
Three watch-list signals need operating proof before escalation
A real movement, or a measurement artifact? The investigator desk says probably both. The more interesting story is not that three signals appeared, but that each one stops just short of becoming an operating conclusion.
The five things · what to leave with
- Tire demand signals are present, but completed RO and utilization data are needed before calling a capacity squeeze .
- Pine Coast Paid Social softness is plausible, but attribution to the channel is not proven .
- Northeast Fleet may be influencing mix, but current evidence could reflect isolated or temporary work .
- Semantic retry downgraded overclaims and preserved the evidence trail.
- Learning memory now carries forward the proof standards for future Snowflake-backed runs.
Editorial standard
Readable like a reporter. Auditable like a workpaper.
Every material statement is paired with an evidence path. Notes and point comments can add color and hypotheses, but they cannot carry numerical claims unless the source data agrees.
The lead
On the surface, the week looks like a manager's early-warning dashboard: tire pressure is rising, Paid Social looks soft in Pine Coast, and Northeast Fleet revenue is moving faster than bookings. But the agent desk found a more useful pattern underneath. Each signal is visible enough to deserve attention, and each is incomplete enough to punish a premature conclusion.
The result is a briefing that deliberately resists drama. The system is not saying the business is in trouble. It is saying the next operating conversation should be narrower, better evidenced, and easier to prove or disprove. That distinction matters: a weak report would convert movement into certainty; this one turns movement into a testable agenda.
Weekly edition
A slower, more synthetic read for the operating week.
The weekly edition is a separate agent run, not a display mode. It widens the story budget and creates its own draft, trace, citations, comments context, and archive record. Instead of treating every signal as a same-day alert, the agents look for pattern durability, repeat evidence, cross-market differences, and whether prior analyst comments changed the story.
Bring forward notes, comments, prior claims, Ask This Report exchanges, and resolved retries.
Ask what changed, what persisted, and what deserves executive attention across the week.
Weekly draft
Not generated in this session yet
Run Weekly Edition to create a separate draft and archive record.
Proof agenda
What the next query has to settle
Tires
Signal: Appointments and capacity index move together .
Missing: Completed ROs, utilization, cancellations, no-shows, and technician availability.
Decision use: Watch staffing and slot availability, but do not call a capacity squeeze yet.
Paid Social
Signal: Reach holds while appointments soften .
Missing: Spend, impressions, CTR, CPL, lead quality, and appointment conversion by source.
Decision use: Inspect funnel efficiency before changing budget or blaming creative fatigue.
Fleet
Signal: Revenue outruns appointment movement .
Missing: Customer concentration, repeat cadence, labor hours, ticket size, and job category mix.
Decision use: Treat as a mix question until repeat work proves durable demand.
Executive Lead
Evidence points to live business questions, not firm conclusions.
The clean takeaway: management has three signals worth tracking, but the file does not yet clear the bar for escalation . The desk should keep them on watch until conversion, utilization, funnel, and repeat-work data close the gap.
That is a useful management posture. It keeps the agenda sharp without overstating the business case. The next meeting should not ask, “Are tires constrained?” It should ask whether completed work, service capacity, and missed demand agree with the appointment signal.
Tires
Demand signal, not a proven capacity constraint.
Repeated tire-related signals suggest momentum . But appointments and revenue can move for reasons that are not operational capacity: price, mix, scheduling behavior, or work that never converts into completed repair orders.
The operational read is therefore cautious. If completed ROs rise with appointment pressure and utilization tightens, this becomes a staffing and slot-management story. If completed work does not move, the signal may be scheduling friction, pricing, seasonal browsing, or demand that is not converting.
Pine Coast Paid Social
Softness is plausible, but attribution remains unproven.
The file supports a possible Paid Social slowdown . It does not prove channel underperformance without spend, impressions, CTR, CPL, source-level conversion, and appointment trend separation.
The story to test is funnel quality, not vague marketing weakness. If spend and reach are stable while click-through and lead-to-appointment conversion fall, creative fatigue becomes plausible. If spend changed, source mapping drifted, or appointment availability tightened, the channel may be taking blame for an upstream or downstream issue.
Northeast Fleet
Possible mix movement, not yet durable.
Fleet revenue may be influencing mix , but the current support could still be a one-off customer, isolated job, or ticket-size effect. Repeat cadence and customer concentration matter here.
The next query should split the movement by customer, job type, and ticket size. A broad fleet shift would show repeated customers and recurring categories. A fragile signal would concentrate in one account, one work order type, or one unusually large ticket.
Why the retry mattered
The first draft wanted a cleaner story than the evidence allowed.
The semantic retry changed the report's center of gravity. The failed version leaned toward declarative claims: tires constrained, Paid Social underperforming, Fleet mix shifting. The accepted version keeps those ideas as hypotheses and names the tests that would make them publishable.
This is the agentic loop doing visible editorial work. It did not merely polish language; it changed claim strength, preserved citations, and left an audit trail for analysts who want to inspect the rejected path.
Point commentary
Comments become feedback for the next agent loop
Comments are saved as point-level context, then classified before they influence future runs. In production they would be written with section ID, paragraph ID, citation IDs, author, timestamp, feedback category, revision intent, and disposition.
Store the comment against the exact claim, citation, chart, paragraph, draft version, and run.
Separate unsupported claims, missing context, tone issues, workflow bugs, preferences, and feature requests.
Turn high-confidence feedback into memory, eval cases, prompt rules, evidence requests, or backlog items.
Run desk
Watch the agentic loop execute
Evidence search
The desk scours, narrows, and preserves the trail
Search queue
Evidence ledger
Hypotheses under test
Retry logic
Where weak drafts are caught before they reach the reader
Rejected overclaim
“Tires are capacity constrained.”
- Missing completed RO proof
- No cancellation/no-show split
- Revenue could be price or mix
Evidence Reviewer sends it back
The retry instruction requires the writer to separate signal detection from causal conclusion, preserve citations, and name the disproof tests.
Rewrite as a watch-list signal. Keep citations. State what would validate or disprove the claim.
Passed editorial gate
“Tire demand signals are present, but completed RO and utilization data are needed before calling a capacity squeeze.”
- Claim strength reduced
- Evidence preserved
- Next checks explicit
Learning mechanism
Run memory becomes habits of success
The system records the causes of retries and successful editorial corrections. Next runs start with those habits loaded into the agent desk.
Lessons carried forward
Next-run behavior changes
- Capacity narratives require completed ROs and utilization before escalation.
- Paid media attribution requires spend, impressions, CTR, CPL, and conversion separation.
- Fleet growth claims require repeat cadence and customer concentration checks.
- Data-quality flags are promoted before any affected metric is used in a headline.
Improvement scorecard
Trace history
Loop history, queries, workpapers, and reasoning summaries
The report can stay clean for readers, while analysts can open the production record behind it: what each agent searched, what it found, what it rejected, and what changed after retry.
Run 01: evidence desk
Retry packet
Rejected draft claim
Tires are capacity constrained; Paid Social is underperforming; Northeast Fleet has shifted mix.
Reviewer reasoning summary
The draft converted weak signals into causal conclusions. The reviewer required completed ROs for tires, source-level funnel metrics for Paid Social, and repeat-customer evidence for Fleet before escalation.
Retry instruction
Rewrite as evidence-bound watch-list signals. Preserve citations. Name what would validate or disprove each claim. Do not use appointment or revenue movement as causal proof.
Learning memory written after run
- Capacity claims must pass completed RO, conversion, cancellation/no-show, and utilization checks.
- Marketing attribution must separate budget, reach, CTR, CPL, and conversion before blaming creative or channel quality.
- Fleet movement must test concentration and repeat cadence before calling a durable mix shift.
- Data-quality blockers are promoted before any affected metric can support the story.
Expandable footnotes
Citations with graphical support
These are the evidence cards the story can lean on. Each figure is drawn with D3 and annotated for the exact point the narrative needs.
[1] Display reporting caveat in Desert ValleyData quality
Display has repeated partial or missing rows, so movement should remain directional and should not become a headline support point.
synthetic_expansion:channel:Desert Valley:Display
0 = ok, 1 = partial, 2 = missing. The annotation highlights why the editor should avoid overclaiming Display revenue.
[4] Tire capacity and appointment pressureService demand
Tire appointments and capacity pressure moved together. The claim still needs completed repair orders before it can become a capacity conclusion.
service_line:Tires:Lakeview:latest_week
The supporting figure shows pressure, not proof. That distinction is the point of the story.
[6] Pine Coast Paid Social funnel softnessMarketing funnel
Impressions can look healthy while appointment yield weakens. The channel story needs source-level funnel metrics before attribution is fair.
channel:Paid Social:Pine Coast:latest_week
The divergence supports investigation, not a final diagnosis.
[8] Northeast Fleet mix movementRevenue mix
Fleet revenue moved faster than bookings, which may imply mix or average-ticket movement rather than broad demand growth.
service_line:Fleet Service:Northeast:latest_week
The desk should inspect RO count, customer concentration, labor hours, and repeat cadence next.
Run archive
Older runs, weekly drafts, and revision history live here
In this demo, run records are saved in browser storage. In production, this page reads from the database: run ID, cadence, prompt, agent trace, citations, draft version, analyst comments, Ask This Report exchanges, retry packets, and publish status.
Analyst notes
Human context that can augment the run
Notes are allowed to shape hypotheses and follow-up checks. The workflow treats them as context, not proof, unless matched to metric evidence.
Deep dive observability
Runtime, handoffs, tools, tokens, retries
Tool calls
load_company_profile, load_channel_rows, load_service_rows, load_analyst_notes, run_deterministic_checks, get_check_results, get_recent_analyst_notes, get_source_summary, compose_briefing
Admin
Schedule routine runs
This hosted demo panel shows the intended operating control: a user can set cadence, instruction, and review mode for scheduled briefings.
About
This runs in production. What follows is the synthetic version of it.
The reporting system on this site is a working demonstration built on invented data for a fictional company. The system it demonstrates is real: a recurring agentic briefing running in production against a national tire and auto service retailer’s performance data, on Snowflake. Everything under In production today below describes that live system. Everything under Designed, not yet built is roadmap, and is labelled that way deliberately — a page where every feature is production is a page you should not trust.
In production today
Snowflake is the evidence layer
Every material number in a published report resolves back to Snowflake. Agents do not freely invent SQL against it. The warehouse is the source of truth for the claims, and the claims are bound to it rather than derived from a summary of it.
The agent harness is the orchestration runtime
Not an IDE assistant — the harness runs as the execution layer, and the agents inside it reach Snowflake through a governed binding maintained by our internal platform. The binding is what makes this safe to run on a schedule: agents get a controlled connection, not credentials and an open SQL prompt.
Cortex Agents and Cortex Analyst are called as tools
The harness calls Cortex Analyst against a semantic layer and Cortex Agents for the analytical work, treating their responses as tool output inside a longer loop it controls. Snowflake owns the question-to-answer step. The harness owns the investigation around it — which question is worth asking, whether the answer contradicts last period, and whether the evidence is strong enough to publish.
Runs execute and log through the OpenAI Agents SDK
Specialist agents, handoffs, tool calls, and traces are captured as first-class records inside an internal application, not as console output. The desk includes Signal Scout, Hypothesis Generator, Investigation Planner, Root Cause Analyst, Evidence Reviewer, Narrative Editor, and Managing Editor.
The loop retries until it clears its bar
Weak drafts are retried rather than published. Missing citations, unsupported numbers, data-quality blockers, and thin causal claims each send the loop back around. It publishes when the success criteria are met, and holding a claim is an acceptable outcome.
The report and its trace publish together
The result is published in an application alongside its logs. The trace is not debugging telemetry filed somewhere else — it ships with the story, because a claim a reader cannot interrogate is a claim they have to take on faith.
What it has actually delivered
Digestion, not discovery
This has not been a discovery engine, and it would be dishonest to sell it as one. It surfaces findings a good analyst would have reached eventually by working through the dashboard. What changed is the digestion: the scanning, the ranking, the evidence assembly, and the defensible writing-up — the work that consumed analyst hours and happened live in meetings — now happens before anyone opens anything.
That is information management, not clairvoyance. A system that reliably produced startling revelations every period would not be investigating; it would be performing, manufacturing a story to justify the run. The bar that makes this trustworthy is the same bar that keeps it unspectacular.
Designed, not yet built
Editable, versioned drafts
Human edits should create new versions, preserve the generated draft, and trigger citation-integrity checks when a user introduces new claims or numbers.
Point comments as context
Comments attached to a specific claim, paragraph, chart, or citation, stored with their target and evidence references, so revision agents can retrieve them as context rather than as unbounded chat.
Feedback as a classified loop
Comments are signals, not instructions. The loop should classify feedback, decide whether it becomes a prompt rule, eval case, evidence request, preference, backlog item, or no action, then expose that disposition in the run history.
Clickable SQL evidence
Superscript markers on material claims opening an evidence drawer: query spec, compiled SQL, source rows, freshness status, chart spec, and the trace events from the agent that used it.
Learning memory across runs
Recording habits of success — proof standards, source reliability, editorial preferences, failed claim patterns — and retrieving them before the next run starts.
Ask This Report
A marginal, grounded conversational layer: highlight a passage, ask a follow-up, get an answer constrained to report context, citations, evidence, notes, and run history — without letting chat replace the crafted narrative.
Full operational layer
Authentication, role-based access, approval workflows, trace redaction, query cost limits, and eval harnesses, so analysts can dive into the machinery beneath an edition-quality report.
Page experience
Edition pages rather than a single scroll. Transitions like paper being laid down: folios update, rules and headlines enter first, evidence unfolds from the margin, the agent timeline moves like a production desk.
What we would want from the platform next
Cortex Analyst answers; it does not remember
Analyst is very good at turning a question into a defensible answer against a semantic view. What it does not do is carry state across invocations — notice that this period’s answer contradicts last period’s, or that a claim it just made was held for thin evidence. We built that loop ourselves. A first-class notion of run continuity would remove the most-copied piece of glue code in systems like this.
Semantic view synonyms are the highest-leverage surface, and the least documented
Users say “non-brand”; the column says IS_BRANDED. Every gap like that is a question Analyst answers worse. Synonyms and comments are not documentation — they are the input the model reasons over, and tuning them beat every prompt change we tried. The docs treat them as a nice-to-have. A worked example with deliberately bad synonyms beside good ones, same question asked of both, would make the case in thirty seconds.
Trace continuity across the boundary
Our orchestration traces live in one system and Cortex’s execution detail lives in another. Stitching a published claim back through the loop and into the query that produced it is work we do by hand. A correlation ID that survives the call into Cortex Agents would make end-to-end provenance close to free — and provenance is the whole product here.
Model names are string literals with no safety net
A wrong model name is a runtime error you hit after writing the query, and the roster moves, so older examples fail on paste. An error that named the available models in the region would eliminate the entire category. The information is clearly present at the point of failure.
Say the grain thing out loud
Classify at the distinct grain, not the row grain. On a real export that is a 20× cost difference, and it is also a correctness issue — at row grain, identical text can receive different labels on different rows and nothing errors. It is a decision made in the first ten minutes and painful to retrofit. It deserves to be in the first tutorial, not learned on a bill.
A reference loop with the guardrails already in it
The gap between “here is how to call a model” and “here is how to safely run what it wrote” is where teams either stall or, worse, do not stall and ship something unbounded. A canonical bounded loop — read-only role, allowlist, row caps, retry limits, stop conditions — would raise the floor across the ecosystem.
Boundary. This public version contains no client data, production credentials, proprietary source systems, private prompts, or live operational workflows. The company, the data, and the analyst notes are invented. The architecture, the workflow, and the editorial discipline are the real ones.
Embedded comments
Reader commentary can sit beside the claims it changes
These are point-specific comments, not analyst notes. Analyst notes inform the agent run broadly; embedded comments are attached to a claim, paragraph, citation, or chart and can be retrieved by revision agents when that exact point is rewritten.