11 min read
Production agents increasingly work differently.
A market-monitoring agent, a news-triage bot or a social-media watcher runs unattended for days or weeks.
The world keeps changing while the agent is idle, and at some point the agent has to decide whether this is the moment to act.
ReLiveGym is a new diagnostic benchmark that targets this gap.
Agents act sparsely over simulated weeks while real-world news, market and social-media streams are replayed in chronological order.
The main finding matters to anyone building long-running agents:
How an agent decides when to act is a harness-design axis in its own right. The best design varies by task, and sometimes by model.
So you can't just pick the strongest model and run it on a 15-minute cron. The trigger policy is part of the system you're evaluating.
Here's what ReLiveGym measures, what has been verified so far, how to get started with the repository, and how to apply its ideas to your own agent harness today.
The paper separates two ideas that are easy to mix up:

Existing long-horizon work doesn't fully capture these challenges, because it mostly assumes an environment that doesn't change over time.
These are exogenously evolving environments: news breaks and prices move whether the agent is awake or not. The agent has to choose when to step in without exhausting a finite token budget.
The same tension shows up in any agent you deploy:

Here is the experimental setup:
| Quantity | Reported value |
|---|---|
| Benchmark tasks | 8 |
| Base language models evaluated | 8 |
| Wake-up (trigger) mechanisms | 3 |
| Continuous-improvement methods | 2 |
| Repeated runs per configuration | 3–6 |
| Episodes per task | 1–3 |
| Paper length | 9 pages (v1, Sept 30, 2026) |

The eight tasks vary along three axes:

The data is real news, market and social-media data, replayed on a simulated clock. The code and data-curation scripts are released, and that the required data is publicly available under permissive licenses.
ReLiveGym evaluates two things together:

That split is the most useful mental model: When a long-lived agent fails, the cause is often the harness (it woke at the wrong time, or never learned from a mistake), not the model's reasoning. For a way to turn harness failures like these into lasting checks, look at turning agent failures into eval tasks.
The harness compares three ways to decide when the agent wakes:
| Trigger | How it works | Typical failure mode |
|---|---|---|
| Sleep | The agent calls a tool to set how long to sleep before its next turn | The agent oversleeps and misses deadlines |
| Watcher | The agent writes a Python script that polls the environment and wakes the LLM only when a condition is met (e.g., a keyword appears in the news) | The condition is wrong or too narrow, so the agent stays asleep through the event that matters |
| Cron | The agent manages a schedule of recurring or one-time triggers | Rigid timing; wakes on schedule whether anything changed or not |

On browser-based tasks, the sleep trigger often failed catastrophically because agents overslept and missed deadlines. Cron acted as a safety net.

That's the "optimal design varies across tasks" result in practice.
The sleep trigger is the cheapest and most flexible, and the watcher is the most efficient when its condition is correct. A dumb fixed schedule can still beat both when a deadline makes a missed wake-up unrecoverable.
The second axis is learning over time.
ReLiveGym tests two continuous-improvement methods that use hindsight feedback, meaning information available after the agent has acted. The goal is to see whether agents can fix the recurring failure modes that show up in long-lived tasks.
Simple test-time feedback mechanisms can significantly reduce persistent strategy failures.

Start by cloning the repository:
git clone https://github.com/SaharaLabsAI/ReLiveGym
cd ReLiveGymThe agent writes or edits a program, does a dry run with validation, and then runs it for real:
write_file / edit_file # author or modify the program
run_program(path, validate=true) # dry-run with validation
run_program(path) # real executionenvkit exists only inside the run_program workspace. It is not a module you can import from the repository. If you try to import envkit in your own scripts, it will fail, and that is expected.
This validate-then-execute step is worth copying even if you never run ReLiveGym. Agent-written programs that run unattended for weeks (watcher scripts are the obvious example) should get a dry run before they're trusted with real wake-ups.

You don't need the benchmark to use its main lesson: make the trigger policy a pluggable, testable part of the harness.
The sketch below shows that. It is my own illustration, not the ReLiveGym API. It implements the three trigger types from the paper, plus a deadline guard that reflects the "cron as safety net" observation, in a form you can replay against logged events.
# Illustrative harness sketch -- NOT ReLiveGym's API.
from dataclasses import dataclass, field
from datetime import datetime, timedelta
from typing import Callable, Optional
@dataclass
class Decision:
wake: bool
reason: str
class SleepTrigger:
"""Agent picks its own sleep duration after each turn."""
def __init__(self):
self.next_wake: Optional[datetime] = None
def set_sleep(self, now: datetime, duration: timedelta):
self.next_wake = now + duration
def check(self, now, observations) -> Decision:
if self.next_wake is None or now >= self.next_wake:
return Decision(True, "sleep elapsed")
return Decision(False, "sleeping")
class WatcherTrigger:
"""Cheap predicate over new observations; LLM wakes only on a match."""
def __init__(self, predicate: Callable[[list], bool]):
self.predicate = predicate
def check(self, now, observations) -> Decision:
try:
hit = self.predicate(observations)
except Exception as e:
# Fail open: a broken watcher should not silence the agent.
return Decision(True, f"watcher error, failing open: {e}")
return Decision(hit, "condition matched" if hit else "no match")
class CronTrigger:
"""Fixed recurring or one-time schedule."""
def __init__(self, times: list[datetime]):
self.pending = sorted(times)
def check(self, now, observations) -> Decision:
if self.pending and now >= self.pending[0]:
while self.pending and now >= self.pending[0]:
self.pending.pop(0)
return Decision(True, "cron fired")
return Decision(False, "not scheduled")
@dataclass
class DeadlineGuard:
"""Wraps any trigger and forces a wake before known deadlines."""
inner: object
deadlines: list[datetime]
margin: timedelta = timedelta(hours=1)
fired: set = field(default_factory=set)
def check(self, now, observations) -> Decision:
for d in self.deadlines:
if d not in self.fired and d - self.margin <= now < d:
self.fired.add(d)
return Decision(True, f"deadline guard for {d.isoformat()}")
return self.inner.check(now, observations)Next, a replay loop. The point is the same as ReLiveGym's: run the same agent against the same replayed history with different trigger policies, and compare the results.
# Illustrative replay loop -- feed it your own logged event history.
def replay(events, trigger, agent_step, start, end, tick=timedelta(minutes=15)):
"""events: list of (timestamp, payload), sorted by time."""
now, i, buffer = start, 0, []
stats = {"wakes": 0, "ticks": 0, "reasons": {}}
while now <= end:
while i < len(events) and events[i][0] <= now:
buffer.append(events[i][1]) # only the past is visible
i += 1
d = trigger.check(now, buffer)
stats["ticks"] += 1
if d.wake:
stats["wakes"] += 1
stats["reasons"][d.reason] = stats["reasons"].get(d.reason, 0) + 1
agent_step(now, buffer) # your LLM call goes here
buffer = []
now += tick
return statsTrack at least three things per policy:
You shouldn't expect a single policy to win on all three across tasks.
You can build an optional cheap shell when check that runs before an agent-prompt cron job is queued:
EIGHT_CRON_WHEN=1 flag.
The stated reason is familiar to anyone running scheduled agents: quiet ticks still cost a full model run. This is a watcher layered over cron, with fail-open semantics so a broken watcher can't quietly turn the agent off. If you want to understand why the sensor layer matters so much here, see how a loop decides whether sensing is cheap or ruinous.
| Evaluation style | Typical focus | What ReLiveGym adds |
|---|---|---|
| Static long-horizon environments | Long action sequences in one persistent state | A world that changes on its own, with sparse wake-ups |
| Short interactive benchmarks | Immediate task completion and tool use | Behavior over simulated weeks, including acting at the right time |
| Web/search agent benchmarks | On-demand retrieval or task completion | Unattended monitoring, recurring actions, temporal relevance |
| Memory / continual-learning evals | Retaining or updating knowledge | Hindsight feedback in a recurring, time-evolving environment, tied to task performance |
It fits alongside related work on aging and lifelong agents:
ReLiveGym complements these rather than replacing them. Its distinctive variables are temporal evolution, sparse intervention, trigger policy and feedback-driven adaptation.
For a single-turn chatbot, quality was mostly about model choice. For an agent that runs for weeks, ReLiveGym shows that when the agent wakes up matters as much.