← The Atlas Journal
Engineering4 September 20269 min read

Grid Agents: the recurring grid work, done for you

A watcher tells you something changed. An agent does the job — the 96-block schedule, the morning brief, the deviation statement — on your schedule, in the format the recipient wants, with a history that says what it produced and what it cost. Here is the architecture, and the check that stops the model inventing a megawatt.

EM
EnergyMap Research Team
India Energy Atlas, CIFR
Illustration by India Energy Atlas Research

The jobs nobody writes about

A large share of every week in Indian power goes on work that is recurring, rule-bound and low-judgement: a 96-block day-ahead schedule for the SLDC and its intraday revisions, a morning market brief for the trading desk, a deviation statement reconciled against SEM meter data, a Monday tracker of orders and tenders from the authorities you follow, a monthly progress report for SECI. Almost all the data behind those jobs is already in the Atlas or in a file the person doing the job already has. What was missing was a worker that does the job on a schedule, in the format the recipient wants, and reports back where they actually look. The watcher we shipped in August was honest about being a watcher: it diffed a source against a cursor and posted an update. It could not produce a spreadsheet, read a file you gave it, run at 07:00 IST on weekdays, email anyone, or tell you what it had done and what it had cost. It was a good scheduler wearing the wrong job title.

What an agent is now

An agent is a definition — a template, your configuration, a five-field cron read in Asia/Kolkata, where the result should go, and a budget — and everything after that is the engine's. Nine stages, and the useful way to read them is as three passes: the clock decides when, the run decides what, and the third pass has to prove the first two happened.

Diagram in three lanes. Top lane, the clock: the definition in agents.definitions carries template, config, a 5-field cron read in Asia/Kolkata, delivery and budget; beat runs every 60 seconds as the only schedule, dispatch-due-grid-agents, with no cron file and no per-agent timer; the claim is FOR UPDATE SKIP LOCKED on due rows and next_run_at moves on in the same transaction. A connector carries one queued row and one task on the grid_agents queue down to the middle lane. Middle lane, one run, deterministic first and the model second: template.pipeline(ctx) fetches and computes every number and each step is a row in agents.run_steps; allow-listed tools are atlas.query, signals.search, files.read_table and sheet.build, all read-only, bounded and budgeted; the validator requires every digit run in the prose to be in the fact ledger or the run fails as 'validator'. A connector labelled checked first — a failed check never writes a file — carries down to the bottom lane. Bottom lane, what comes back, where a run is not succeeded until all three are true: artifacts as xlsx, csv, md or pdf written from the pipeline's tables and never from the model, kept 90 days; delivery to inbox, email, Slack webhook or download, each one a receipt row, and any failure is partial; history as steps, tokens, cost and every receipt on the run page and in agents.usage_monthly.
Fig. 1 — The nine stages. Names are from app/grid_agents/: engine.py, store.py, artifacts.py, delivery.py and runtime/validators.py.

The clock is the one thing we kept unchanged from the watcher, because it was already right. Beat runs a single task every sixty seconds; the claim is one statement with FOR UPDATE SKIP LOCKED that advances next_run_at in the same transaction, so a due agent is dispatched at most once per window without a distributed lock anywhere. What changed is what happens after the claim: instead of an evaluator returning drafts, a run gets a budget, a working directory, a cancel flag and a step ledger, and it can write files.

The rule that decides arguments in this codebase is the third pass. A run is succeeded only when the artifact exists and every delivery the user asked for returned 2xx or was explicitly skipped. If the spreadsheet was written but the Slack post failed, that is Partial, and the run page says so with the receipt attached. If the model's prose failed the validator, that is Failed with the validator's message, not a prettier guess. Silent success is the worst failure this feature could have, and every status in the vocabulary exists to avoid it.

Three minutes to a morning brief

Creation is a five-step flow with a live preview card, and the job list is the whole product surface in one screen. Four templates ship built to work — Morning market brief, Day-ahead RE schedule (96 blocks), DSM reconciliation and Regulatory tracker — plus a blank agent you describe in your own words, and the four legacy watchers, which survive as quick Watch jobs. What a given account can pick is decided on the server; a card the account cannot start is drawn greyed and still readable rather than hidden.

Step one of the Grid Agents creation flow, captured on a free account. A five-step stepper reads Job, Configure, Schedule, Report, Review. Nine template cards are laid out in a grid: Morning market brief, Day-ahead RE schedule (96 blocks), DSM reconciliation, Regulatory tracker, Dataset update watch, Demand threshold, Tender watch, Policy watch and Blank agent. Four of the nine are greyed with a small lock or clock glyph and the word Soon, and remain fully readable. A live preview card on the right shows the new agent's name and its four delivery glyphs. A Continue button sits below the grid.
Fig. 2 — Step 1 on a free account. No price and no tier name appears anywhere in the console; a locked card is a lock glyph, and the only upgrade surface is the sheet that links to pricing.

The form for step 2 is not written by hand. It is rendered from the template's own JSON Schema, which is why adding a template is a backend pull request and nothing else — the same property the watcher wizard had, carried forward and extended with file, location and instructions field kinds so a DSM agent can take your schedule and SEM exports and an RE agent can take a plant coordinate.

The last step is the one that matters. Before Create agent is enabled you run a test, synchronously, on half the per-run budget, with every delivery channel except download switched off. You see the real timeline, the real first rows of the real spreadsheet, and the real narrative — and if the test fails, the button stays disabled. Nobody should find out on Monday at 07:00 that their column mapping was wrong.

Step five of the creation flow, Review, after a successful test run. The agent is named National morning brief. A Run a test button sits beside a green Succeeded chip and the summary sentence: yesterday's DAM averaged 4,512 rupees per MWh across 96 blocks, the evening peak block cleared highest. Below it a six-step timeline: Read yesterday's DAM 3 s, Read demand and RE share 3 s, Compared to last Thursday 2 s, Wrote the brief 10 s, Output checked 1 s, Built the sheet 2 s. Below the timeline a file row for brief_national_2026-09-04.xlsx at 29 kB with a three-row preview table of block, DAM rupees per MWh and RTM rupees per MWh, marked First rows only. Back and Create agent buttons sit at the bottom.
Fig. 3 — The test run is the same code path as a scheduled run, with trigger='test'. There is no separate preview engine to drift.

After that, /agents is one shell with three levels — the dashboard, an agent, a run — and moving between them changes the level of the same page rather than navigating. The run view is the audit: a timeline of every step with the number it produced, the token split, the artifacts with an in-page preview of the first sheet, and a receipt line per delivery channel with the time it went out.

The run level of the Grid Agents console. Breadcrumb reads Grid Agents, Karnataka morning brief, Schedule 4 Sept R1. A green Succeeded chip and 09:40 · 37 s sit at the top right with re-run, copy-link and report glyphs. The timeline on the left has five steps: Fetched 96 blocks 4 s, Built 96 blocks · 289 MWh 11 s with a source disclosure, Wrote the brief 10 s, Wrote schedule.xlsx 5 s, Sent to inbox 7 s. The right column shows a stats card reading Tokens 11k, Of month 0.37 percent, Tools 4, then three artifact cards — schedule_2026-09-04_R1.xlsx at 25 kB with Blocks, Notes and Assumptions sheet tabs and a scrollable preview table, blocks.csv at 25 kB, and brief.md at 25 kB rendering a Morning brief heading — followed by Kept 90 d and two delivery receipts, Inbox at 09:41 and desk@example.in at 09:41.
Fig. 4 — The run level. Every step title carries its own number, and every delivery carries the time it was accepted.

Deterministic first, the model second

A template is code that fetches and computes. The model — GLM 5.2, through OpenRouter, one call per template run — writes a paragraph about what the code found. It never produces a cell in the spreadsheet. That is a design, and on its own a design is just an intention, so it is enforced by a check.

The check is deliberately blunt: every maximal run of digits in the model's prose must appear in the run's fact ledger, which is populated from the data path itself — the rows the tools returned and the numbers the pipeline computed. It looks at digits rather than claims, so it cannot be talked around by rephrasing. Years, ordinals and the small counts a sentence is entitled to use are allow-listed, because firing on “the first of three states” would only teach the model to avoid writing readable English. Paste this into a REPL:

app/grid_agents/runtime/validators.py — the check, and a run of it against the fact ledger from a production run
import re

_NUMBER = re.compile(r"\d[\d,]*(?:\.\d+)?")
ALWAYS_ALLOWED = frozenset(
    {str(n) for n in range(0, 25)} | {str(year) for year in range(2000, 2101)}
)


def normalise(token: str) -> str:
    """'4,512.00' -> '4512'. The comparison form for both sides."""
    text = token.replace(",", "").strip().lstrip("+-")
    if "." in text:
        text = text.rstrip("0").rstrip(".")
    text = text.lstrip("0") if text.strip("0") else "0"
    return text or "0"


def unsupported_numbers(text: str, facts: set[str]) -> list[str]:
    """The numbers in 'text' that the fact ledger cannot account for."""
    offenders: list[str] = []
    seen: set[str] = set()
    for match in _NUMBER.finditer(text or ""):
        token = normalise(match.group(0))
        if token in ALWAYS_ALLOWED or token in facts or token in seen:
            continue
        seen.add(token)
        offenders.append(match.group(0))
    return offenders


# The facts the MORNING_BRIEF pipeline handed the model on 2026-09-04.
facts = {normalise(str(v)) for v in (5048.13, 6162.75, 254138, 26.9)}

print(unsupported_numbers(
    "DAM averaged ₹5,048.13/MWh. RTM averaged ₹6,162.75/MWh. "
    "Peak demand fell to 254,138 MW and RE share to 26.9%.", facts))
print(unsupported_numbers(
    "DAM averaged ₹5,048.13/MWh, and evening peak touched 254,900 MW.", facts))

# []
# ['254,900']

The second sentence is the failure mode this exists for. It is fluent, it is plausible, it is within a rounding error of the truth, and it is a number no code in the system ever produced. The run fails with error_code='validator' and the offending token in the message. Crucially it is not one of the three consecutive failures that pause an agent: a bad narrative is a fact about one run's output, not a broken configuration, so the agent keeps its schedule and the user keeps being told.

Two things the rule taught us. Comparing magnitudes rather than signed values is the honest scope of a digit-run check — an interchange flow stored as -2164 was unmatchable, and a narrative that correctly wrote “2,164 MW from Uttar Pradesh to Haryana”, with the direction in words the way a flow is actually written, was being rejected as a hallucination. And a response that satisfies the schema while saying nothing passes the digit check trivially, because it has no numbers to be wrong; roughly one narrative call in three came back as three dots before the system prompt forbade it, so a placeholder rule sits beside the digit rule.

What a run costs, in tokens

Every model call writes its input, output and reasoning tokens to the run row, and those roll up into a monthly ledger you can read. The unit the product speaks in is tokens: an agent's share of a month is a number and a percentage, and nothing about model spend is rendered as currency in the console.

The Grid Agents usage page for September 2026, showing 41 runs. A tokens card reads 412k of 3.0M with a segmented progress bar split between three agents — Karnataka morning brief 300k, Tumkur 50 MW day-ahead 100k and CERC plus KERC weekly 12k — and the line Resets 1 Oct IST. Below it a tokens-per-day bar chart over 14 days totalling 255k tokens, with daily bars between roughly 9k and 26k. Below that a by-agent table with runs, tokens and share columns: Karnataka morning brief 22 runs, 300k, 73 percent; Tumkur 50 MW day-ahead 15 runs, 100k, 24 percent; CERC plus KERC weekly 4 runs, 12k, 2.9 percent. No currency appears anywhere on the page.
Fig. 5 — The allowance, in tokens. The month resets on the 1st, IST.

The numbers are modest for a template, because a template run makes exactly one model call and the rest is code. The MORNING_BRIEF run in our production acceptance set finished in 34.4 s across nine steps — seven fetches, one narrative call, one validation — using 2,411 input, 2,656 output and 1,509 reasoning tokens, and produced three artifacts. A blank agent doing free-form work over live grid data is a different animal: one CUSTOM run in the same set walked into its 120,000-token ceiling and stopped, which is the budget guard doing its job.

Two details in how the cap is enforced. It stops a run being claimed rather than killing one in flight, so nobody loses half a spreadsheet to an allowance boundary; the agent is paused with a reason and told when the month resets. And a run that hit its own ceiling is kept out of the reliability numbers entirely — a plan limit working as designed is not an outage, and folding it into a success-rate alert would only make the alert measure our users' ambition.

What is next

The roadmap templates are visible in the gallery with real config schemas rather than a waiting list: the CEA daily thermal return, outage and tripping intimations, standalone intraday revisions, open-access day-ahead requests, monthly developer progress reports, energy-account confirmation and REC application packs. Portal submission is deliberately not among them — agents prepare a filing in the right format, validated and revision-numbered; a human still presses submit.

We also ran the “agents that need agents” spike, and the answer was to build rather than adopt. Both candidates — Hermes Agent self-hosted, and DigitalOcean Gradient's managed agents — can be made to write their steps into our tables, which was the decision rule we had set in advance. Neither can enforce a budget or a tool allow-list per step, which is the thing that actually protects a user from a runaway job, and which our own bounded loop already does. When v2 needs an agent to spawn one, it will be a second bounded loop in our own engine writing into the same run_steps table.

Grid Agents is live at /agents; the gallery, including the roadmap, is at /agents/templates.

PRs (private repositories): BE schema, schedules, entitlements, gateway · BE run engine, tools, artifacts, delivery · Morning brief + regulatory tracker · 96-block schedule + DSM reconciliation · FE console shell · FE creation flow · FE agent and run levels · FE usage, gallery, inbox cutover · Droplet, worker, metrics, runbook · Sub-agent decision memo · product page on Confluence

Filed under
Engineering, published 4 September 2026
← More from The Atlas Journal