← The Atlas Journal
Engineering31 August 202616 min read

My Grid Agents: watchers over India's grid

A watchlist is a list you have to remember to look at. An agent is a template, your configuration, a frequency and a worker that checks for you. Here is the contract that stops it telling you the same thing twice, and the scheduler that is one SQL statement.

EM
EnergyMap Research Team
India Energy Atlas, CIFR
Illustration by India Energy Atlas Research

Here is a thing an analyst covering Indian power actually wants: every day, check whether any new solar tender has appeared in Rajasthan, and only bother me about the large ones. It is a small ask. It is also, for most people, an unpaid job — a browser tab, a calendar reminder, a folder of PDFs, and a nagging suspicion that something was missed in the week they were on leave.

Atlas has had a “My Grid” watchlist for a while. It is a list of states, datasets and trackers a reader has said they care about, and it does exactly what a list does: it sits there. The reader still has to go and look. This release turns that relationship around. A Grid Agent is a template, a configuration you fill in, a frequency, and a worker that runs on that schedule and writes back what changed. The reader stops checking, and starts being told.

The whole feature is about four hundred lines of evaluator, one Celery task, one SQL statement and four tables. What made it hard was not any of that. It was the promise underneath — that an agent tells you a thing once — and this post is mostly about how that promise is kept.

From a list to a watcher

An agent has four parts, and it is worth being precise about which piece owns what, because the correctness argument later depends on it.

The template is code: a named watcher with a Pydantic config model, a floor on how often its source can move, and an evaluate function. The config is the reader's answer to that model — the states, the keywords, the threshold. The frequency is how often they want it checked. And the cursor is the evaluator's own memory of where it got to, which the reader never sees and never sets.

The evaluator is a pure function of those first three plus a clock. It reads its source, it returns what to publish and where it got to, and it writes nothing. Everything that touches a database on the way out is the engine's job.

services/tools-api/app/home/agents/base.py — the contract a template must satisfy
class Evaluator(Protocol):
    key: str
    title: str
    description: str
    config_model: type[BaseModel]
    min_frequency: AgentFreq
    min_plan: Plan

    def evaluate(
        self, cfg: BaseModel, cursor: dict, now: datetime
    ) -> tuple[list[UpdateDraft], dict]:
        """Return the drafts to publish and the cursor to store.

        Must not write to anything. Must be idempotent: calling it again with
        the cursor it just returned yields ([], same_cursor).
        """
        ...

There is one more rule that is easy to miss and matters a great deal in practice. The first run of a brand-new agent primes: it records where the world is and publishes nothing. An agent answers “what has changed since you created me”, not “what does the archive contain”. Without that, the first thing a reader ever sees from their first agent is a wall of history they did not ask for, and they turn it off.

The four things an agent can watch

Four templates ship in v1. Each one names a source and a cursor shape, and the cursor shape is the interesting column — it is what determines how hard idempotency is for that template.

TemplateWatchesSourceCursor
DATASET_UPDATE_WATCHData Shop datasets you follow, for a refreshtools.shop_products.updated_at{dataset_id: last_seen_at}
DEMAND_THRESHOLDMetered demand for a region or state crossing a level you setatlas_intelligence.electricity_demand{last_breach_at}
TENDER_WATCHNew tenders, filtered by state and technologySignals ledger, tender family{last_event_id}
POLICY_WATCHNew regulatory documents from issuers and on topics you chooseSignals ledger, institutional family{last_event_id}

Three more are on the roadmap and are not built: a substation-commissioned watcher, a data-centre pipeline watcher, and a price alert. They are listed in the wizard with real config schemas and marked coming_soon, which is a deliberate middle position — hiding them makes the roadmap invisible, and enabling them would promise a watcher that does not exist. A create request naming one is refused by the same code path that refuses a template key that was never real, because from a write path's point of view “not built yet” and “does not exist” are the same answer.

Two of the four have a template floor coarser than the plan's and two do not, which is worth stating plainly because it decides what the wizard offers. The frequency a reader may actually pick is the coarser of the template's own floor — how often its source can possibly move — and their plan's floor. A free reader can run one agent and check it daily; Atlas Pro raises both. That is the only sentence about plans in this post, and the details live on the pricing page.

The cursor is the whole promise

A worker queue redelivers. A container restarts mid-task. A deploy replays a window. Any of those, in a naive design, means a reader gets told about the same tender twice, and two duplicate alerts is roughly the amount of noise it takes for someone to stop trusting a notification stream entirely.

So the property is stated as a property, not as a hope: fed the cursor it just returned, an evaluator must produce zero drafts. Each of the four has its own test for that, rather than one generic loop that could pass by accident.

app/home/tests/test_agents_evaluators.py — prime, report, then repeat with the cursor it returned
def test_dataset_watch_primes_then_reports_then_is_idempotent(db, dataset_ids):
    evaluator = DatasetUpdateWatch()
    cfg = DatasetUpdateConfig(dataset_ids=dataset_ids)

    # First run records where the datasets are and says nothing: a new agent
    # must not announce refreshes that predate it.
    drafts, cursor = evaluator.evaluate(cfg, {}, NOW)
    assert drafts == []
    assert set(cursor) == set(dataset_ids)

    with db.cursor() as cur:
        cur.execute(
            "UPDATE tools.shop_products SET updated_at = %s WHERE slug = %s",
            (NOW, dataset_ids[0]),
        )
    db.commit()

    drafts, cursor = evaluator.evaluate(cfg, cursor, NOW)
    assert len(drafts) == 1
    assert drafts[0].payload["datasets"][0]["dataset_id"] == dataset_ids[0]

    # Idempotency: the cursor it just returned yields nothing new.
    repeat, repeat_cursor = evaluator.evaluate(cfg, cursor, NOW)
    assert repeat == []
    assert repeat_cursor == cursor

For the two ledger templates this is nearly free, and the reason is the shape of the source rather than anything clever in the evaluator. The Signals ledger is an append-only log with a monotonic event_id. Diffing it is a keyset read — everything with an id greater than the one I last saw, ordered by id, capped at a page — and the cursor is that id. The same cursor always selects the same rows, so a replay is not a hazard; it is a re-read that finds nothing.

app/home/agents/evaluators.py — TenderWatch, on the cursor it owns
class TenderWatch:
    """New tenders in the Signals ledger, filtered by state and technology.

    Cursor is {last_event_id} — the ledger's own identity column. Event-log
    semantics make the diff trivial and a replay safe: the same cursor always
    selects the same rows.
    """

One detail in there is not obvious and took a bug to find. The cursor advances past everything the query read, including rows a config filter then dropped. Holding it back to the last row the reader actually saw feels tidier and is wrong: those filtered rows would be re-read on every run for ever, and the page cap means a busy ledger could starve the agent of the rows it does want. What the reader is shown and where the evaluator got to are two different facts.

The other reason the ledger shape matters is that adding a template on top of it is almost no work. POLICY_WATCH is the same keyset read as TENDER_WATCH with a different document family and different filters. A new signal family becomes a new template with no new machinery, which is the return on consuming an event log rather than polling a mutable table.

Why a threshold needs a different rule

DEMAND_THRESHOLD gets none of that. There is no event to have an id: the question is “is the peak in my window above my line right now”, and the answer is a fresh reading every time. Storing the last reading and diffing it would give you an alert on every run, because the number always moves a little.

What a reader means by “tell me when northern-region evening peak goes above sixty-five gigawatts” is: tell me when it starts. A peak that sits above the line for a week is one event, not seven. So the cursor is a latch. It is set while a breach is live and cleared on the first reading back the right side of the line, and the alert fires only on the set. This is the re-arm rule, and it is the single most important line of behaviour in the template.

app/home/agents/evaluators.py — the four lines that make a sustained breach one alert
        if not breaching:
            return [], {"last_breach_at": None}
        if cursor.get("last_breach_at"):
            return [], dict(cursor)
Line chart of five successive runs of a demand-threshold agent against a 65,000 MW threshold drawn as a dashed line. Run 1 reads 60,000 MW, below the line, and the cursor is null — armed. Run 2 reads 70,000 MW, crosses the line, and an alert fires; the cursor's last_breach_at is set. Run 3 reads 71,000 MW, still above, and the agent is silent because the cursor is already set. Run 4 reads 40,000 MW, a recovery, and the cursor is cleared back to null — re-armed. Run 5 reads 80,000 MW and alerts again. Under each reading a chip shows whether last_breach_at is set or null after that run.
Fig. 1 — The five readings are the sequence asserted by test_demand_threshold_alerts_once_and_rearms_only_after_recovery. Two alerts across five runs, three of which are above the line.

There is a companion rule that looks like a special case and is really the same one. If the query returns no reading at all — a gap in the feed, a collector down — the cursor is left exactly as it is. It is tempting to treat “no data” as “not breaching” and clear the latch, and that would re-alert on the first row to arrive after an outage, for a breach the reader was already told about. No reading is not a recovery.

pytest -v — the two properties, on a real Postgres carrying the source tables
app/home/tests/test_agents_evaluators.py::test_demand_threshold_alerts_once_and_rearms_only_after_recovery PASSED
app/home/tests/test_agents_evaluators.py::test_demand_threshold_is_idempotent_on_its_own_cursor PASSED
app/home/tests/test_agents_evaluators.py::test_demand_threshold_holds_a_breach_through_a_gap_in_the_feed PASSED
app/home/tests/test_agents_evaluators.py::test_dataset_watch_primes_then_reports_then_is_idempotent PASSED
app/home/tests/test_agents_evaluators.py::test_tender_watch_primes_reports_and_is_idempotent PASSED
app/home/tests/test_agents_evaluators.py::test_policy_watch_reports_by_issuer_and_topic_then_is_idempotent PASSED

============================== 31 passed in 4.73s ==============================

Claiming the agents that are due

Every agent carries a next_run_at. The scheduler's job is to find the ones whose time has come, dispatch each exactly once for that window, and move their clocks forward. The obvious failure mode is two dispatchers — a rolling deploy, a second beat container someone left running — dispatching the same agent for the same window and producing two runs.

The usual answers are a distributed lock in Redis, a broker feature, or a single-instance constraint enforced by hope. We used none of them. The claim and the clock advance are one SQL statement, so they are one transaction, and Postgres does the rest.

app/home/agents/store.py — sync_claim_due_agents, run by beat every 60 seconds
WITH due AS (
    SELECT id, frequency
      FROM home.grid_agents
     WHERE is_active AND next_run_at <= NOW()
     ORDER BY next_run_at
       FOR UPDATE SKIP LOCKED
     LIMIT %s
)
UPDATE home.grid_agents a
   SET next_run_at = NOW() + CASE due.frequency
           WHEN '15m' THEN INTERVAL '15 minutes'
           WHEN '1h'  THEN INTERVAL '1 hour'
           WHEN '6h'  THEN INTERVAL '6 hours'
           WHEN '24h' THEN INTERVAL '24 hours'
           WHEN '7d'  THEN INTERVAL '7 days'
       END
  FROM due
 WHERE a.id = due.id
RETURNING a.id

Two choices in there are load-bearing. SKIP LOCKED rather than a plain FOR UPDATE, because a second dispatcher should move on to other work rather than queue behind the first and then dispatch a window that has already been dispatched. And the UPDATE in the same statement as the SELECT, because the instant the first dispatcher commits, those rows are no longer due — so even after the locks are released, the second dispatcher claims nothing.

That is a claim worth testing rather than asserting, so there is a test that opens two real connections and interleaves them.

app/home/tests/test_agents_engine.py — two dispatchers, one due agent
def test_concurrent_dispatchers_claim_an_agent_at_most_once(db, owner):
    agent_id = _make_agent(db, owner, frequency="1h")

    with psycopg.connect(_dsn()) as a, psycopg.connect(_dsn()) as b:
        with a.cursor() as cur_a:
            claimed_a = store.sync_claim_due_agents(cur_a)
        assert agent_id in claimed_a          # A holds the lock, uncommitted

        with b.cursor() as cur_b:
            claimed_b = store.sync_claim_due_agents(cur_b)
        assert agent_id not in claimed_b      # skipped, not blocked

        a.commit()

        with b.cursor() as cur_b:
            claimed_after = store.sync_claim_due_agents(cur_b)
        assert agent_id not in claimed_after  # no longer due
        b.commit()

The batch is capped at two hundred agents per claim, so a backlog drains over several ticks instead of one tick holding locks across the whole table. The transaction commits before anything is enqueued, because the claim's only purpose is to make those rows not-due for everyone else, and a lock held open while Redis is being talked to lasts as long as the broker is slow.

The beat schedule for the entire feature is one entry, at sixty seconds. There is no cron file anywhere, and no per-agent timer — the clock lives in a column.

Diagram in three lanes. Top lane, every 60 seconds: beat runs the only schedule, dispatch-due-agents; the claim is FOR UPDATE SKIP LOCKED with a limit of 200 over active and due rows; the same statement advances next_run_at to now plus the frequency in the same transaction; one run_agent task is enqueued per claimed id after the lock is released. Middle lane, one run: load config from grid_agents and the cursor from agent_state then close the snapshot; evaluate config, cursor and now, reading the source read-only and returning updates and a new cursor; one commit writes agent_runs, agent_updates and the new cursor together or not at all; the rail shows an unread badge and the drawer marks it read. Bottom lane, when it raises: the evaluator throws, the cursor is untouched and the run is written as error so the next window re-reads the same range, three consecutive errors are counted off the agent_runs log, and the owner is told in an alert update and the agent is paused.
Fig. 2 — Dispatch, one run, and the failure path. Names are from app/home/agents/store.py, engine.py and base.py.

One more decision that a reader feels immediately: a new agent is created with next_run_at = now(). It runs on the next tick rather than one full frequency later, so someone who sets up a daily agent does not spend a day wondering whether it works. On a clean stack the first run landed just over two seconds after the create call, against a sixty-second tick.

make e2e-agents-lifecycle — dispatched by the real beat container, not by the test process
1. Create a tender agent — the first run is immediate
  [PASS] next_run_at is now, not now+frequency — 2026-08-30T21:41:14.024774Z
  [PASS] the agent_state row exists with an empty cursor — {}

2. Waiting up to 150s for the beat container to claim it and the worker to prime it
  [PASS] a run row appeared without this process dispatching anything — 1 run(s) after 2.1s
  [PASS] the priming run reports nothing — no_change
  [PASS] the cursor now records where the ledger is — {'last_event_id': 315}
    beat latency: 2.1s (tick = 60s)

3. A new tender in the ledger -> beat dispatch -> update in the rail
  [PASS] the second run was dispatched by beat — 2 run(s) after 60.0s
  [PASS] the rail badge shows 1 unread — {"total_unread":1}
  [PASS] opening the drawer marks it read — {"marked":1}
  [PASS] and the unread count drops to zero — {"total_unread":0}

What a broken agent should do

Sources go down. A column gets renamed. A config that resolved fine last month stops resolving. The question is not whether an evaluator will raise, but what the system owes the reader when it does.

The first half of the answer is the cursor rule, and it falls straight out of the contract. A run that raises is recorded as an error, and the cursor is left exactly where it was. The next window therefore re-reads the same range, so a transient outage costs a delay and nothing else. The engine commits the run row, every update and the new cursor together — an update the reader can see and a cursor saying it has been seen must become true at the same instant, or a crash between them either loses an update for ever or repeats it every window.

app/home/agents/engine.py — the cursor moves only on a run that did not raise
        with connection.cursor() as cur:
            run_id = store.sync_record_run(...)
            update_ids = [
                store.sync_insert_update(cur, ...) for draft in drafts
            ]
            # The cursor moves only on a run that did not raise. This is the
            # retry guarantee: a failed run re-reads its range next window.
            if error is None:
                store.sync_set_cursor(cur, agent_id, new_cursor)
            store.sync_mark_ran(cur, agent_id)

        # One commit: the run row, every update and the cursor land together.
        connection.commit()

The second half is the part we argued about. Retrying for ever is the safe-looking option, and it is the wrong one. An agent that fails silently for ever is worse than no agent, because the reader believes they are covered. Three consecutive errors are not the weather; they are the configuration. So the strike count is read off the run log, and on the third the agent publishes an alert update its owner can actually see and pauses itself.

app/home/agents/engine.py — three strikes
# Consecutive failures before an agent is paused.
ERROR_STRIKES = 3

PAUSE_TITLE = "This agent is failing; check its configuration"

A success between failures resets the count, which is what stops a flaky upstream from slowly pausing everything a reader owns. And an unknown template key fails the run rather than the worker — a bad row in one agent must not take down the process that is running everyone else's.

pytest -v — the failure behaviour, asserted rather than described
app/home/tests/test_agents_engine.py::test_run_writes_the_run_updates_and_cursor_together PASSED
app/home/tests/test_agents_engine.py::test_a_failed_run_records_the_error_and_leaves_the_cursor_untouched PASSED
app/home/tests/test_agents_engine.py::test_three_consecutive_errors_alert_the_user_and_pause_the_agent PASSED
app/home/tests/test_agents_engine.py::test_a_success_between_failures_resets_the_strike_count PASSED
app/home/tests/test_agents_engine.py::test_running_a_deleted_agent_is_not_an_error PASSED
app/home/tests/test_agents_engine.py::test_an_unknown_template_key_fails_the_run_rather_than_the_worker PASSED

There is one optional piece that is worth naming precisely because it is the sort of thing a release post usually oversells. An update's summary can be rewritten by a small language model into two sentences. That flag is called AGENT_LLM_SUMMARY, it defaults to false in settings, the shipped .env.example sets it to 0, and this release ships with it off. What readers see today is the deterministic string the evaluator wrote, which is also what makes a screenshot reproducible and a test assertable. When it is switched on, every failure mode — timeout, non-200, malformed JSON, empty completion, no endpoint configured — returns the template string unchanged, because a formatting nicety must never take down the run that produced it.

One schema, and no frontend work

This is the design decision that decides whether the feature grows or ossifies, and it is worth spelling out because the tempting alternative is very easy to write.

A wizard that configures four different templates wants, badly, to be a switch on the template key with a hand-built form per branch. It reads fine, it ships fast, and it means every new template is a backend change plus a frontend change plus a release to line them up. Two templates later nobody adds templates any more.

So the wizard never learns the template keys. Each template's config model is serialised with Pydantic's model_json_schema() and served on GET /agents/templates, and the form is rendered from that and nothing else.

model_json_schema() for DEMAND_THRESHOLD, exactly as the templates endpoint serves it
{
  "properties": {
    "region": {
      "description": "Regional grid code (NR, WR, SR, ER, NER), IN for all-India, or a state name.",
      "title": "Region",
      "type": "string"
    },
    "metric": {
      "default": "peak_demand_mw",
      "enum": ["peak_demand_mw", "latest_demand_mw"],
      "title": "Metric",
      "type": "string"
    },
    "op": { "default": ">", "enum": [">", "<"], "title": "Op", "type": "string" },
    "value": {
      "description": "Threshold in MW.",
      "title": "Value",
      "type": "number"
    },
    "window_hours": {
      "default": 24,
      "description": "How far back the peak is taken. Ignored for latest_demand_mw.",
      "maximum": 168,
      "minimum": 1,
      "title": "Window Hours",
      "type": "integer"
    }
  },
  "required": ["region", "value"],
  "title": "DemandThresholdConfig",
  "type": "object"
}

The mapping is small and completely mechanical. An enum becomes a select. An array of enum becomes multi-select chips; an array of string becomes the same control with free entry. A number becomes a numeric input, with its unit read out of the schema's own description — that is where Threshold in MW. puts MW next to the box. A boolean becomes a checkbox, a string becomes a text field. Anything the parser cannot classify becomes a text field carrying the raw JSON rather than vanishing, because a silently dropped field is a configuration the reader believes they set.

The claim “adding a template is backend-only work” is only worth making if something enforces it. Two kinds of test do. One feeds the parser a schema this build has never seen and asserts it produces a usable form, seeds a config from the schema's defaults alone, and validates against its constraints. The other reads every file in the agents UI and asserts none of them contains a template key.

npx vitest run lib/home/agent-schema.test.ts — the enforcement, not the intention
 ✓ parseConfigSchema — a template this build has never seen > classifies every kind the handoff names, including array-of-enum
 ✓ parseConfigSchema — a template this build has never seen > follows $ref into $defs for the enum members
 ✓ parseConfigSchema — a template this build has never seen > reads the unit out of the new schema's own description
 ✓ parseConfigSchema — a template this build has never seen > seeds a config from the schema's defaults alone
 ✓ parseConfigSchema — a template this build has never seen > validates against the new schema's own constraints
 ✓ no per-template frontend code > lib/home/agent-schema.ts names no template key
 ✓ no per-template frontend code > components/home/agents/SchemaFields.tsx names no template key
 ✓ no per-template frontend code > components/home/agents/NewAgentWizard.tsx names no template key
 ✓ no per-template frontend code > components/home/agents/AgentEditPanel.tsx names no template key
 ✓ no per-template frontend code > components/home/agents/AgentsRail.tsx names no template key
 ✓ no per-template frontend code > components/home/agents/AgentListItem.tsx names no template key
 ✓ no per-template frontend code > components/home/agents/AgentUpdatesDrawer.tsx names no template key
 ✓ no per-template frontend code > lib/home/agents.ts names no template key

 Test Files  1 passed (1)
      Tests  31 passed (31)

The cost is real and worth naming. A schema-driven form is never as nice as one somebody designed by hand for a particular question, and there are two places where a bespoke control would be better. What we get back is that a template is a pull request against one Python file, and the wizard picks it up on the next deploy of a service it does not share a release with. Given how many templates this feature ought to have in a year, that is the trade to make.

Run the whole thing on your laptop

None of this needs a cloud account. The stack is Postgres, Redis, the API, a Celery worker and beat, all under Docker Compose. What follows was run on a --depth 1 clone at the Phase 4 tip, on a Mac, with the ports overridden because other checkouts already held the defaults.

A clean clone, the stack up, and the schema applied
$ git clone --depth 1 \
    https://github.com/India-Energy-Atlas/espresso-india-transmission-map.git blog-clean-p4
$ cd blog-clean-p4 && git log --oneline -1
8b9c14b feat(home): agent lifecycle e2e under the real beat, demo seed proof [IEA-2588] (#1284)

$ export ATLAS_DB_PORT=5464 ATLAS_REDIS_PORT=6389 ATLAS_API_PORT=8020 \
         COMPOSE_PROJECT_NAME=blogp4
$ export DATABASE_URL="postgresql://grid:grid@localhost:5464/grid" \
         REDIS_URL="redis://localhost:6389/0" API_BASE="http://localhost:8020"

$ make up
 Container blogp4-db-1  Healthy
 Container blogp4-redis-1  Healthy
 Container blogp4-worker-1  Started
 Container blogp4-beat-1  Started
 Container blogp4-api-1  Started
waiting for api health...
stack up: api http://localhost:8020

$ make migrate
...
applying sql/migrations/tools_049_my_grid_agents.sql
applying sql/migrations/tools_055_home_phase4.sql
migrations applied

make demo loads the metric fixtures and seeds the sample agents — one on the free demo reader, four on the pro one, each with a pre-generated update, so every screen has real content in it rather than an empty state.

make demo, then make verify-demo reading it back
$ make demo
seeded 13 tools, 3 news items
  demand rows            504
  frequency samples      864
  fuel mix + carbon rows 4032
  registry rows          88
  tracker ledger rows    102
  regulatory documents   184
metric fixtures loaded
seeded demo users + {'tool_usage_events': 90, 'dataset_view_events': 37,
                     'pins': 4, 'agents': 5, 'agent_updates': 5}

$ make verify-demo
5 seeded agent(s) across the demo users:
  demo_home_free   Followed datasets refreshed      DATASET_UPDATE_WATCH   24h  active=True updates=1
                   -> [info] 1 dataset refreshed
  demo_home_pro    CERC and RERC storage orders     POLICY_WATCH           24h  active=True updates=1
                   -> [info] 3 new regulatory documents from CERC, RERC
  demo_home_pro    Followed datasets refreshed      DATASET_UPDATE_WATCH   24h  active=True updates=1
                   -> [info] 2 datasets refreshed
  demo_home_pro    NR evening peak above 65 GW      DEMAND_THRESHOLD       24h  active=True updates=1
                   -> [alert] NR demand above 65,000 MW
  demo_home_pro    Rajasthan solar tenders          TENDER_WATCH           24h  active=True updates=1
                   -> [notable] 2 new solar tenders in Rajasthan

ALL ASSERTIONS PASSED

From there, open /home, create an agent of your own in the wizard, and watch the rail. Or run the whole lifecycle unattended. The lifecycle target deliberately never calls the dispatcher itself: every run in it is claimed by the beat container and evaluated by the worker container, because a stubbed scheduler is exactly the part most likely to be broken in a way unit tests miss. It takes about five minutes, most of which is waiting for real sixty-second ticks.

The full lifecycle, on the same clean clone
$ time make e2e-agents-lifecycle
...
6. Edit the config — the cursor resets and the next real run re-evaluates from scratch
  [PASS] the cursor was thrown away with the question it answered — {}
  [PASS] beat ran the edited agent — 4 run(s) after 59.9s
  [PASS] the re-primed run reports nothing rather than replaying history — no_change
  [PASS] no update was duplicated by the reset — 2 update(s), was 2
  [PASS] and the cursor was re-established from scratch — {'last_event_id': 317}

7. Delete — runs, updates and state go with the agent
    before delete: runs=4 updates=2 state=1
    after delete:  runs=0 updates=0 state=0 agent=0
  [PASS] every child row cascaded — (0, 0, 0, 0)

ALL ASSERTIONS PASSED

make e2e-agents-lifecycle  1.50s user 0.32s system 0% cpu 5:03.73 total

Step six there is the cursor rule seen from the other end. Editing an agent's config throws its cursor away, because the old cursor answers a question the reader no longer asked. The next run primes again — and publishes nothing, rather than replaying the history the previous config had already covered.

What this version leaves open

Delivery is in-app only. An agent update lands in the rail and the drawer; the email path is modelled in the delivery column and not wired. The three roadmap templates are roadmap. The summariser is off. And the whole engine currently reads four sources, which is a small number for a product whose interesting future is a reader assembling their own watchers over everything Atlas tracks.

The thing we would defend hardest, though, is the boring part. Not the templates — templates are cheap now. It is that (config, cursor, now) signature, and the insistence that feeding an evaluator its own output produces nothing. Every other property in this post is downstream of it: the retry safety, the crash safety, the concurrent-dispatcher safety, the fact that an edit can reset state without duplicating anything. Get that contract right and a scheduler is one SQL statement. Get it wrong and no amount of locking will save you.

The underlying numbers these agents watch are all public. National and regional demand is on /forecast, exchange prices are on /iex-market, and the datasets a DATASET_UPDATE_WATCH agent follows are listed on /data. The same series with a state focus sit on the state pages — Rajasthan, Gujarat and Tamil Nadu among them — and an agent is, in the end, just a standing question asked against those.

Sources: services/tools-api/app/home/agents/base.py, engine.py, store.py, evaluators.py, registry.py, llm.py, services/tools-api/app/home/celery_app.py, sql/migrations/tools_055_home_phase4.sql, lib/home/agent-schema.ts and components/home/agents/NewAgentWizard.tsx in the India Energy Atlas repositories. Every terminal, source and test capture on this page is pasted from a run on a clean --depth 1 clone at the Phase 4 tip, on 30 August 2026; both figures are rendered by scripts in scripts/blog/.

Filed under
Engineering, published 31 August 2026
← More from The Atlas Journal