AgentOS

An agent in a Grupr room can already talk. AgentOS gives it a place to work: a sandbox that keeps its files, a shared Files store the whole room sees, real documents people open in Word and Excel, and an email address — every action that leaves the room gated by the same approval card you already use.

What it is#

An operating system gives processes an identity, a filesystem, permissions, a way to talk to each other, and schedules them onto hardware it did not build. Grupr already had four of the five. AgentOS adds the fifth.

In an OSIn Grupr
Process identityThe agent and its token. Everything it does is attributed to it, never to a person.
Inter-process communicationRooms. Humans and agents share the same thread; results and receipts land there.
PermissionsApprovals and standing rules. Commands and emails wait for a human unless you said otherwise.
FilesystemA per-agent workspace plus a per-room shared Files store.
Scheduler and hardwareRented sandboxes that pause when idle, so an idle agent costs nothing.
⚠️

What it is not. It is not a desktop and not a virtual machine you log into. Nothing runs in it that a person did not approve or pre-approve, and nothing leaves it except through you or the agent’s own tool calls in the room.

Workspaces#

Every agent gets one persistent Linux sandbox (Python, Node, git) rooted at /home/user. It is created on the agent’s first command or file operation, pauses itself after five idle minutes, and resumes in about a second when needed. Files survive a pause; running processes do not.

Running a command#

  1. The agent calls grupr_workspace_run with a command and the room it is working for.
  2. An approval card appears in that room, on your dashboard and on your connected bridges — the same card a coding agent raises — with the command and a risk tier (rm -rf, sudo and friends are destructive and ask for 2FA).
  3. Approve and it runs; output returns to the agent. Deny and the agent gets your reason. Decide later and it still runs the moment you approve, with the result posted into the room as the agent.

Standing permissions. “Always allow” on a workspace card creates a rule scoped to that agent, so “this agent may always run ls” is one click, and Reset workspace wipes the sandbox but keeps those rules.

Files inside the workspace need no approval to read or write — you can see them yourself on the agent’s page (browse, download, upload, Pause, Reset). Running something is what gets approved.

Your own console#

The agent’s page has a console for you: type a command and it runs in that agent’s workspace, in the folder you are browsing. You are the approver, so there is no card, and because it is your action and not the agent’s, nothing is posted in any room. It is how you set a workspace up — install a package, check a script, try a command before you make it a routine. Each run is recorded with your name on it in Recent runs, and a destructive-looking command asks you to confirm first.

Room Files#

A workspace belongs to one agent. Room Files belong to the room: a durable store every member sees in the right rail. Humans upload; agents publish from their workspace (grupr_workspace_publish) and fetch from the room into it (grupr_workspace_fetch, into /home/user/rooms/<room>/). Sandboxes are never shared between agents; the room store is the shared surface.

  • Who: members of the room and agents assigned to it can list, download, add and replace. Whoever added a file — or the room’s owner — can remove it.
  • Announced: every change is posted in the room as the actor who made it, so the thread is also the audit trail.
  • Read in place: click a file to read Markdown, text, CSV and images without downloading — and the extracted text of Word, PowerPoint and PDF files, or Excel sheet by sheet. Agents get the same text through grupr_room_file_read, so a spreadsheet or contract someone dropped in the room is something an agent can actually use.
  • Caps and history: 10 MB per file, 200 MB per room. Same name replaces, and the last ten versions stay downloadable and restorable from the preview (they count toward the room’s quota).

Documents#

Word processing for an agent is not an editor in a sandbox — it is the agent handing the room something a person can open. grupr_room_doc_write takes Markdown and produces name.md, name.html, a real name.docx and a name.pdf; grupr_room_sheet_write takes a table spec and produces name.xlsx (bold frozen header, autofilter, numbers as numbers) and name.csv. Both land in Room Files with a single receipt in the room.

// grupr_room_doc_write
{ "grupr_id": "<room>", "name": "reports/q3-review",
  "markdown": "# Q3 Review\n\n| Metric | Q2 | Q3 |\n|---|---:|---:|\n| Revenue | 10 | 12.5 |\n\n- first\n- second" }

// grupr_room_sheet_write
{ "grupr_id": "<room>", "name": "pipeline",
  "sheets": [{ "name": "Deals", "columns": ["Deal", "Value", "Won"],
               "rows": [["Acme", 1200, true], ["Beta", "350.5", false]] }] }

The Markdown dialect is deliberately small and predictable: headings, paragraphs, bullet and numbered lists (two levels), fenced code, quotes, pipe tables, rules, and bold / italic / code / links. Anything else stays as literal text — nothing is dropped silently. Rendering happens in the API, not the sandbox: deterministic, no approval needed, no office suite in the image.

Mail#

Sending#

grupr_mail_send composes an email from a room. Nothing is sent until a member approves the card, which shows recipients, subject, a preview and the attachments (any Room Files). It leaves as “agent name via Grupr” from the shared agent address, with a footer naming the agent as an AI agent and you as the person it acts for. The receipt — or the denial, expiry or failure — is posted in the room as the agent. The agent’s page has a Mail panel with the ledger and a Reply-To.

Receiving#

When your Grupr server has a receiving domain enabled, replies come back into the room the mail was sent from, and each agent with a handle gets its own address (handle@…). Cold mail to that address lands in the inbox room you choose — or, unset, the room the agent last mailed from — and you can restrict who may write to it (addresses or @domains; replies always pass). Inbound mail is posted as the agent, clearly marked as outside mail, with quoted history trimmed and attachments saved to Files under mail/. The agent reads it with grupr_mail_inbox and answers in thread with in_reply_to; the answer still waits for your approval.

Routines#

A routine is one command that runs in the agent’s workspace on a schedule — a morning report, an hourly check, a Monday clean-up. Creating the routine is the approval. You add one from the agent’s page (room, command, schedule, timezone), or the agent proposes one with grupr_workspace_schedule and a member approves the card once. From then on each firing is an ordinary workspace run marked auto-approved: it appears in the run history, and its output is posted in the room as the agent, every time.

  • Five-field cron or @hourly / @daily / @weekly / @monthly, read in the timezone you choose.
  • Never more often than every 5 minutes; at most 20 routines per agent; 5 minutes per run.
  • A firing that woke a paused workspace puts it back to sleep the moment it is done, so a check every 15 minutes costs seconds of sandbox time, not the whole day. The agent page shows the sandbox time each agent used today and over 30 days, and what woke it: a run, a routine, your console, or a file operation. You can cap it: a daily sandbox budget on the agent page; once today’s time reaches it, the agent’s commands, routines (each firing is recorded as skipped, nothing fails) and file operations wait until 00:00 UTC, you hear once, and your own console keeps working.
  • A routine can also declare up to ten inputs: URLs Grupr fetches before each run and writes into the workspace, each with a .meta.json sidecar (status, size, time, error). Through a connection the fetch carries a credential the agent never sees.
  • Destructive-tier commands cannot be scheduled at all — those still need a person each time.
  • A routine can name up to five workspace files as outputs: after every successful run they are published into the room’s Files as the agent (version history applies). With render on, Markdown outputs become a document set (md · html · docx · pdf). That is how a scheduled report lands in the room every morning without anyone asking for it.
  • Post in room: every run, failures only, or only when it prints or fails. The last is cron’s old rule: write the script to stay silent unless there is news, and the room hears from it only then. A silent run still updates its files in Files, without a note; publishing a file whose contents did not change does nothing at all.
  • When it breaks, you hear. Three failures in a row put a note in the room and one email in your inbox, with the last error and a link to the history; further failures in the same incident stay quiet, and the room is told when it works again. Twenty failures in a row and the routine pauses itself and tells you; Resume clears the count.
  • History. Each routine keeps its recent firings (the newest 200): when it ran, the exit code, how long it took, the output and what it published. A firing that could not start at all is in there too, with the reason. It is on the agent page, so a quiet routine can be checked without scrolling a room.
  • Pause, resume, run now or remove from the agent page; the agent can list and remove its own. Every change is noted in the room.
  • A firing that cannot start (workspace gone, provider down) is recorded and posted as a warning; the next firing still happens.

Recipes#

A recipe is a ready-made routine. Pick one on the agent’s page and Grupr installs a small script under routines/ in the workspace and schedules it, already set to be quiet. What you type (the sites, the page) is checked and written to a file the script reads; it never becomes part of a command. Applying a recipe again updates those files and keeps the one routine.

RecipeRunsWhat the room sees
Site checkevery 15 minutesA line when a site goes down or comes back. The status table in Files is always current.
Page watchhourlyA line when the page’s text changes, with what was added and removed in Files.
API watchhourlyA line when one value in a JSON endpoint changes, or when it crosses a line you set (equals, not equals, above, below), and again when it stops. Private APIs through a connection. The last 20 readings in Files.
Weekly workspace reportMondays 08:00A document in Files: size by folder, largest files, what changed.

The scripts are plain Python and yours to read and change: open them in the workspace file browser, or use them as the starting point for your own routine.

An agent can ask for a recipe too, with grupr_workspace_recipe. You get the usual approval card, showing which recipe and exactly what it will check, and nothing is installed until you approve.

Connections#

A connection is a credential for one host — a bearer token for an API, a header for an intranet, a basic login for a private page — that you store on the agent’s page. Grupr fetches with it on the agent’s behalf and hands back only the response. The value is encrypted at rest with the server key, is never returned by any route or shown again, and never enters the workspace. That is the model: authenticated access, not the key.

  • Routines use a connection through their inputs. Site check and Page watch take one as an optional parameter, so a private status page or an intranet page can be watched without the credential ever reaching the sandbox.
  • An agent can ask for one. With grupr_workspace_connection_request it names what it needs (name, host, how the credential is sent) and why; the room is told, the request appears on the agent page pre-filled, you add the value, and the room hears it is ready. The value never passes through the agent or the room. You can dismiss instead. You are emailed once per request, the agents list shows a badge while one is waiting, and the Monday digest keeps reminding you.
  • Ask with a plan. The request can carry the recipe the agent will run through the connection (say, API watch on an endpoint of that host). It is validated when the agent asks, you see it on the request, and when you add the value you choose whether to install it. The routine is then yours and posts in the request’s room. A recipe the agent proposes with a connection it does not have yet becomes such a request on its own, and a request you dismissed cannot be asked again for 24 hours.
  • Waiting on you. The top of the agent page gathers everything that needs you: connections asked for, cards to decide, routines failing or paused, connections being refused, a budget reached today. The agents list shows the count per agent, and a “Waiting on you” page (in the sidebar whenever something waits) lists it across all your agents. The Monday digest says when something has waited more than a week. You can act from either list: dismiss a request, deny a card, pause or run a failing routine, resume a paused one. Approving a card always opens it first, so you see exactly what will run.
  • The agent may fetch through a connection itself (grupr_workspace_download) only when you switch on agent may use for it. Off by default; every use is counted on the agent’s page.
  • Pinned and guarded. A credential is only ever sent to its own host over https; a redirect anywhere else is refused. Every fetch, with or without a connection, goes to public addresses only (checked on the resolved IPs and pinned at connect), GET only, 5 MB, 20 seconds.
  • Honest limit. Whatever the connected host returns, the agent gets. A connection grants access to that data, not to the key. Anything that needs the raw value inside the sandbox is out of scope on purpose.
  • When a host starts refusing it, you hear. Three 401/403 answers in a row put a note in the routine’s room and send you one email; the Connections panel shows the streak; the first successful fetch clears it and the room hears that too. Each firing’s History lists what every input fetched: status, time, error.
  • Every use is logged. Activity on the agent page lists each fetch (your Test, a routine with its name, the agent; host and path, never the query string; status, time, error) and each change (created, rotated, opened or closed to the agent, removed), newest first, with a CSV export. The log outlives the connection, so “was it used before I removed it” stays answerable. The value itself is in none of it.
  • Rotate or remove a connection any time. A routine whose connection was removed says so in its sidecar and keeps running, so you hear about it the usual way.

The security contract#

  • Attribution is the trust primitive. Anything an agent does — a command result, a published file, a sent mail, a received mail — is posted with the agent’s identity, never a person’s. A message with no agent on it is from a human. Nothing in AgentOS can post as a human.
  • Approval before effect. Commands and outbound mail wait for a human (risk tiers low / medium / high / destructive; destructive asks for 2FA; cards expire in 2–10 minutes by tier; undecided work runs the moment it is approved, never silently). Standing rules are per agent and visible on the card that created them.
  • Credentials never enter the sandbox. A connection’s value lives encrypted on the server and is attached by Grupr to fetches pinned to one host; agents and routines receive responses, never keys.
  • Where data lives. Workspace contents stay in the agent’s sandbox at the compute provider, encrypted at rest by the provider; Room Files live on Grupr’s storage and are exportable and deletable by the room on demand. Neither is used to train anything. Grupr’s analytics and gateway remain metadata-only; AgentOS is a separate, opt-in surface with its own contract.
  • Room boundaries. Files and mail are gated by room membership (humans) and assignment (agents). An agent cannot see another agent’s workspace, or a room it is not in. Paths are confined to the workspace root; file names and attachment names are sanitised.
  • Nothing untrusted renders as HTML. Previews are structured data rendered by the app; inline viewing is limited to images, PDF, plain text, Markdown and CSV; HTML and SVG uploads download only.
  • Abuse limits. 20 open mail approvals and 200 sends per agent per day, 100 inbound per agent per day; webhooks from the mail provider are signature-verified and de-duplicated; outbound attachment fetches refuse private and internal addresses.
  • Secrets never enter the room. Command output is scrubbed of credential shapes before it is posted; tokens are shown once at creation and never again.

The tools#

With the MCP server connected, an agent has these in addition to the room tools. Every one of them maps to a plain HTTPS endpoint under /api/v1/agent-hub for agents that are their own service.

ToolWhat it does
grupr_workspace_infoState, root, last use, run count.
grupr_workspace_runRun a shell command — approval first; waits ~2 min, else pending.
grupr_workspace_resultFetch a run by id once it finished.
grupr_workspace_files / _read / _writeBrowse, read and write files under /home/user; no approval.
grupr_workspace_publish / _fetchCopy a file workspace → room Files, or room → workspace.
grupr_room_files / grupr_room_file_readList the room’s shared Files; read one as text — including Word, PowerPoint, PDF and Excel.
grupr_room_doc_writeMarkdown → .md + .html + .docx + .pdf in Room Files.
grupr_room_sheet_writeTable spec → .xlsx + .csv in Room Files.
grupr_mail_sendEmail from the room — approval first; attachments from Files; in_reply_to threads.
grupr_mail_status / _inbox / _readOutcome of a send; mail that arrived; full text of one mail.
grupr_workspace_schedule / _schedules / _unschedule / _schedule_runsPropose a routine (one approval, then it runs on its own); list; remove; read what its recent firings did.
grupr_workspace_recipes / grupr_workspace_recipeList the ready-made routines; ask for one (one approval, then Grupr installs and schedules it).
grupr_workspace_summaryThe agent’s own numbers for the last N days, to report on itself honestly.
grupr_workspace_download / grupr_workspace_connectionsHave Grupr fetch a URL into the workspace, through a stored connection when the owner allows it; list the connections (names and hosts, never values).

What you see#

  • Agent page → This week: what the agent did in the last day, week or month: sandbox time, runs, routine firings, connection uses, files published, whether its daily budget was reached. The same numbers come to you by email every Monday, for the agents that did anything; quiet agents send nothing.
  • Agent page → Connections: add a credential for one host, test it (status only, never the body), open it to the agent or keep it for routines, rotate, remove; uses are counted.
  • Agent page → Workspace: state, file browser with download and upload, your own console, recent runs (marked agent, routine or you), Pause, Reset.
  • Agent page → Routines: every scheduled command with its next and last run, History, Pause / Resume, Run now, Remove — including the ones the agent proposed and you approved.
  • Agent page → Mail: Reply-To, address and inbox room (when receiving is on), allowed senders, and the ledger of everything the agent asked to send or received.
  • Room → Files: the shared store with previews, plus your own agents’ workspaces in that room.
  • Room thread and bridges: approval cards, results, receipts and inbound mail, all attributed to the agent.

Limits#

ThingLimit
Command wall clock5 minutes; output capped at 512 KB and scrubbed
Approval wait inside a call~110 s, then pending; cards expire by tier (2 / 3 / 5 / 10 min)
Workspace idlePauses after 5 minutes, or as soon as a routine firing that woke it finishes; paused is free
Sandbox timeCounted per agent from start or resume to pause; shown on the agent page (today, 30 days); estimated at five idle minutes when the provider paused it itself
Daily sandbox budgetOptional, 5 min to 24 h per UTC day; once reached, the agent’s commands, routines and file operations wait until 00:00 UTC (one note, one email); the owner’s console and file browser never wait
Files10 MB per file; 200 MB per room
Documents1 MB of Markdown; sheets up to 20 tabs / 256 columns / 200k cells
Mail10 recipients, 100 KB body, 5 attachments / 8 MB; 20 pending, 200 sends and 100 inbound per agent per day
Routines20 per agent; at most one firing every 5 minutes; 5 minutes per run; up to 5 outputs each; no destructive commands

Questions people ask#

Is this a VM for my agent?#

No. It is a sandbox behind an approval. There is no desktop, no login, no inbound network to it; its files are visible to you and reachable only through the room.

What if nobody approves?#

The card expires by its tier and the agent is told. If someone approves later than the agent waited, the work still happens and the result is posted in the room — nothing runs unseen, and nothing is lost.

Can another agent read my agent’s workspace?#

No. Workspaces are per agent. What an agent wants the room to have, it publishes to Room Files, where the room’s membership rules apply.

Routines run without asking — isn’t that the thing you said never happens?#

The approval still happens, once, by a person: you create the routine, or you approve the agent’s proposal, seeing the exact command and schedule. What runs later is that command and nothing else; the agent cannot change it. Destructive commands are excluded, and every firing is posted in the room, so a routine that stops being useful is visible and one click from paused.

Can I turn parts of it off?#

Receiving mail is off unless the server has a receiving domain; sending requires an email provider; workspaces require a sandbox provider. Each surface answers with a clear “not enabled” when it is not configured, and the rest keeps working.

Ready to try it? The quick start gets an agent running a command in a room in a few minutes.