Grid Agents: the recurring grid work, done for you
A watcher tells you something changed. An agent does the job — the 96-block schedule, the morning brief, the deviation statement — on your schedule, in the format the recipient wants, with a history that says what it produced and what it cost. Here is the architecture, and the check that stops the model inventing a megawatt.
The jobs nobody writes about
A large share of every week in Indian power goes on work that is recurring, rule-bound and low-judgement: a 96-block day-ahead schedule for the SLDC and its intraday revisions, a morning market brief for the trading desk, a deviation statement reconciled against SEM meter data, a Monday tracker of orders and tenders from the authorities you follow, a monthly progress report for SECI. Almost all the data behind those jobs is already in the Atlas or in a file the person doing the job already has. What was missing was a worker that does the job on a schedule, in the format the recipient wants, and reports back where they actually look. The watcher we shipped in August was honest about being a watcher: it diffed a source against a cursor and posted an update. It could not produce a spreadsheet, read a file you gave it, run at 07:00 IST on weekdays, email anyone, or tell you what it had done and what it had cost. It was a good scheduler wearing the wrong job title.
What an agent is now
An agent is a definition — a template, your configuration, a five-field cron read in Asia/Kolkata, where the result should go, and a budget — and everything after that is the engine's. Nine stages, and the useful way to read them is as three passes: the clock decides when, the run decides what, and the third pass has to prove the first two happened.

app/grid_agents/: engine.py, store.py, artifacts.py, delivery.py and runtime/validators.py.The clock is the one thing we kept unchanged from the watcher, because it was already right. Beat runs a single task every sixty seconds; the claim is one statement with FOR UPDATE SKIP LOCKED that advances next_run_at in the same transaction, so a due agent is dispatched at most once per window without a distributed lock anywhere. What changed is what happens after the claim: instead of an evaluator returning drafts, a run gets a budget, a working directory, a cancel flag and a step ledger, and it can write files.
The rule that decides arguments in this codebase is the third pass. A run is succeeded only when the artifact exists and every delivery the user asked for returned 2xx or was explicitly skipped. If the spreadsheet was written but the Slack post failed, that is Partial, and the run page says so with the receipt attached. If the model's prose failed the validator, that is Failed with the validator's message, not a prettier guess. Silent success is the worst failure this feature could have, and every status in the vocabulary exists to avoid it.
Three minutes to a morning brief
Creation is a five-step flow with a live preview card, and the job list is the whole product surface in one screen. Four templates ship built to work — Morning market brief, Day-ahead RE schedule (96 blocks), DSM reconciliation and Regulatory tracker — plus a blank agent you describe in your own words, and the four legacy watchers, which survive as quick Watch jobs. What a given account can pick is decided on the server; a card the account cannot start is drawn greyed and still readable rather than hidden.

The form for step 2 is not written by hand. It is rendered from the template's own JSON Schema, which is why adding a template is a backend pull request and nothing else — the same property the watcher wizard had, carried forward and extended with file, location and instructions field kinds so a DSM agent can take your schedule and SEM exports and an RE agent can take a plant coordinate.
The last step is the one that matters. Before Create agent is enabled you run a test, synchronously, on half the per-run budget, with every delivery channel except download switched off. You see the real timeline, the real first rows of the real spreadsheet, and the real narrative — and if the test fails, the button stays disabled. Nobody should find out on Monday at 07:00 that their column mapping was wrong.

trigger='test'. There is no separate preview engine to drift.After that, /agents is one shell with three levels — the dashboard, an agent, a run — and moving between them changes the level of the same page rather than navigating. The run view is the audit: a timeline of every step with the number it produced, the token split, the artifacts with an in-page preview of the first sheet, and a receipt line per delivery channel with the time it went out.

Deterministic first, the model second
A template is code that fetches and computes. The model — GLM 5.2, through OpenRouter, one call per template run — writes a paragraph about what the code found. It never produces a cell in the spreadsheet. That is a design, and on its own a design is just an intention, so it is enforced by a check.
The check is deliberately blunt: every maximal run of digits in the model's prose must appear in the run's fact ledger, which is populated from the data path itself — the rows the tools returned and the numbers the pipeline computed. It looks at digits rather than claims, so it cannot be talked around by rephrasing. Years, ordinals and the small counts a sentence is entitled to use are allow-listed, because firing on “the first of three states” would only teach the model to avoid writing readable English. Paste this into a REPL:
import re
_NUMBER = re.compile(r"\d[\d,]*(?:\.\d+)?")
ALWAYS_ALLOWED = frozenset(
{str(n) for n in range(0, 25)} | {str(year) for year in range(2000, 2101)}
)
def normalise(token: str) -> str:
"""'4,512.00' -> '4512'. The comparison form for both sides."""
text = token.replace(",", "").strip().lstrip("+-")
if "." in text:
text = text.rstrip("0").rstrip(".")
text = text.lstrip("0") if text.strip("0") else "0"
return text or "0"
def unsupported_numbers(text: str, facts: set[str]) -> list[str]:
"""The numbers in 'text' that the fact ledger cannot account for."""
offenders: list[str] = []
seen: set[str] = set()
for match in _NUMBER.finditer(text or ""):
token = normalise(match.group(0))
if token in ALWAYS_ALLOWED or token in facts or token in seen:
continue
seen.add(token)
offenders.append(match.group(0))
return offenders
# The facts the MORNING_BRIEF pipeline handed the model on 2026-09-04.
facts = {normalise(str(v)) for v in (5048.13, 6162.75, 254138, 26.9)}
print(unsupported_numbers(
"DAM averaged ₹5,048.13/MWh. RTM averaged ₹6,162.75/MWh. "
"Peak demand fell to 254,138 MW and RE share to 26.9%.", facts))
print(unsupported_numbers(
"DAM averaged ₹5,048.13/MWh, and evening peak touched 254,900 MW.", facts))
# []
# ['254,900']The second sentence is the failure mode this exists for. It is fluent, it is plausible, it is within a rounding error of the truth, and it is a number no code in the system ever produced. The run fails with error_code='validator' and the offending token in the message. Crucially it is not one of the three consecutive failures that pause an agent: a bad narrative is a fact about one run's output, not a broken configuration, so the agent keeps its schedule and the user keeps being told.
Two things the rule taught us. Comparing magnitudes rather than signed values is the honest scope of a digit-run check — an interchange flow stored as -2164 was unmatchable, and a narrative that correctly wrote “2,164 MW from Uttar Pradesh to Haryana”, with the direction in words the way a flow is actually written, was being rejected as a hallucination. And a response that satisfies the schema while saying nothing passes the digit check trivially, because it has no numbers to be wrong; roughly one narrative call in three came back as three dots before the system prompt forbade it, so a placeholder rule sits beside the digit rule.
What a run costs, in tokens
Every model call writes its input, output and reasoning tokens to the run row, and those roll up into a monthly ledger you can read. The unit the product speaks in is tokens: an agent's share of a month is a number and a percentage, and nothing about model spend is rendered as currency in the console.

The numbers are modest for a template, because a template run makes exactly one model call and the rest is code. The MORNING_BRIEF run in our production acceptance set finished in 34.4 s across nine steps — seven fetches, one narrative call, one validation — using 2,411 input, 2,656 output and 1,509 reasoning tokens, and produced three artifacts. A blank agent doing free-form work over live grid data is a different animal: one CUSTOM run in the same set walked into its 120,000-token ceiling and stopped, which is the budget guard doing its job.
Two details in how the cap is enforced. It stops a run being claimed rather than killing one in flight, so nobody loses half a spreadsheet to an allowance boundary; the agent is paused with a reason and told when the month resets. And a run that hit its own ceiling is kept out of the reliability numbers entirely — a plan limit working as designed is not an outage, and folding it into a success-rate alert would only make the alert measure our users' ambition.
What is next
The roadmap templates are visible in the gallery with real config schemas rather than a waiting list: the CEA daily thermal return, outage and tripping intimations, standalone intraday revisions, open-access day-ahead requests, monthly developer progress reports, energy-account confirmation and REC application packs. Portal submission is deliberately not among them — agents prepare a filing in the right format, validated and revision-numbered; a human still presses submit.
We also ran the “agents that need agents” spike, and the answer was to build rather than adopt. Both candidates — Hermes Agent self-hosted, and DigitalOcean Gradient's managed agents — can be made to write their steps into our tables, which was the decision rule we had set in advance. Neither can enforce a budget or a tool allow-list per step, which is the thing that actually protects a user from a runaway job, and which our own bounded loop already does. When v2 needs an agent to spawn one, it will be a second bounded loop in our own engine writing into the same run_steps table.
Grid Agents is live at /agents; the gallery, including the roadmap, is at /agents/templates.
PRs (private repositories): BE schema, schedules, entitlements, gateway · BE run engine, tools, artifacts, delivery · Morning brief + regulatory tracker · 96-block schedule + DSM reconciliation · FE console shell · FE creation flow · FE agent and run levels · FE usage, gallery, inbox cutover · Droplet, worker, metrics, runbook · Sub-agent decision memo · product page on Confluence