Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -220,7 +220,7 @@ The edges are documented, not denied — the full list with mechanisms is in the
- The in-process sandbox is defense in depth; `ContainerExecutor` is the real boundary. `run_command` children are unconfined.
- The HTTP API does not yet use the durable session layer.
- On the Claude CLI backend an agent node is *delegated*, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the `bypass` tier — explicit opt-in — has no checks at all.
- Policy documents govern planning; the tool plane still reads CLI flags.
- Agent tool policies are enforced on tool-calling backends. A policy document refuses delegated Claude CLI runs; allowed tool calls are not written to the document audit.
- The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version `0.1.8` · [changelog](CHANGELOG.md) · [roadmap](ROADMAP.md) · [website](https://codegraphcontext.github.io/GraphARC/) · MIT
18 changes: 12 additions & 6 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -324,9 +324,11 @@ The component with no prior art to copy. It exists, and the cycle runs.
runner from claiming a session, and nothing reclaims one whose runner died
holding it. That is a claim, not a lease.

## 7. Policy engine — `[~] ~75% built, 0% wired`
## 7. Policy engine — `[~] node, edge and tool paths wired`

Everything here works and nothing calls it.
The planner's admission gate and the agent CLI compile policy documents into
their existing permission boundaries. Compiled rules do not audit ordinary
allowances; the agent records document-driven denials and approval decisions.

- [x] **7.1 — Declarative policy config** over nodes, edges, tools and spend.
TOML in, `PolicyEngine` out; a commented example ships at
Expand All @@ -339,7 +341,7 @@ Everything here works and nothing calls it.
`engine.approval_router(handlers, tenant=…)` produces the callback a
`Harness` already obeys, and `engine.permission_policy(tenant=…)` produces
a real `PermissionPolicy`.
- [x] **7.3 — Policy versioning and decision audit.** Every decision lands in a
- [x] **7.3 — Policy versioning and decision audit.** Every direct engine check lands in a
JSONL record naming the resource, subject, tenant, effect, the rule id and
reason that produced it, the policy version, and a digest of the document
— so a decision can be tied to the exact policy text that made it.
Expand All @@ -355,9 +357,13 @@ Everything here works and nothing calls it.
longer imported by nothing. What the compiled object still cannot carry is
what `permission_policy()` cannot either: the approver role and the audit
record, because `EdgePolicy.decide` returns a bare `Decision`. Admission
treats `ask` as not-yet-permitted. Still open: no call from `AgentNode` or
`grapharc agent` to `permission_policy()`, so the tool plane is still
governed by Python objects rather than by the document.
treats `ask` as not-yet-permitted. `grapharc agent --policy` now compiles
tool rules through `permission_policy()`, pairs document `ask` with
`approval_router()`, and records document denials through `check_tool()`.
Flags can narrow the document but cannot widen it. Policy and tenant
follow the shared flag/environment/config precedence, and a configured
policy refuses delegated execution just as an explicit one does. Ordinary
allowed tool calls still have no document-audit record.

## 8. Memory & artifacts — `[~] ~85%`

Expand Down
37 changes: 37 additions & 0 deletions docs/cookbook/03-agents-and-tools.md
Original file line number Diff line number Diff line change
Expand Up @@ -1601,3 +1601,40 @@ grapharc agent --workspace ./scratch --deny 'run_command' --ask 'write_file' \

A trace lands at `<workspace>/trace.jsonl` either way; `grapharc trace`,
`grapharc metrics` and `grapharc viz` read it.

### How does a policy document govern the CLI agent?

`grapharc agent --policy policy.toml --tenant acme` applies the document's
`resource = "tool"` rules through the same harness that filters tool schemas
before the model sees them. Flags can narrow that authority: any denial wins,
then any approval requirement, and a tool runs freely only when both sides
allow it. `--allow '*'` cannot override a document denial.

Policy and tenant can also come from configuration:

```toml
# grapharc.toml
[grapharc]
policy = "policy.toml"
tenant = "acme"
```

The document must declare `acme` when it lists tenants. Resolution is
`flag > GRAPHARC_POLICY / GRAPHARC_TENANT > grapharc.toml > default`.
`--config PATH` selects another file; a relative policy path in that file is
anchored to its directory. Parent directories are not searched. This agent
config path resolves policy and tenant only; model, budgets and the other agent
options retain their CLI/Python-argument defaults.

A governed JSON result includes `config_file`, `sources`, `policy_source`, and
`policy_document` with the document's path, source, version, digest and tenant.
The human view includes the document's source beside its path. The
`policy_audit` path points to `policy-audit.jsonl` beside the trace, recording
document-driven denials and approval decisions. Ordinary allowed calls and
flag-only denials are not document-audited; attempted calls and their outcomes
remain on the run trace.

Document approval requests fail closed under `--json` or redirected stdin.
A policy selected from flags, environment or config also refuses
`--executor claude-cli` and Claude CLI models: that delegated loop cannot
enforce the document. Bad config is refused before model setup or delegation.
6 changes: 6 additions & 0 deletions docs/cookbook/07-slack.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,6 +116,12 @@ Configuration is environment-only, read once at startup:
| `GRAPHARC_SLACK_LIVE_INTERVAL` | `2.5` | seconds between two edits of the status message |
| `GRAPHARC_SLACK_LIVE_URL` | unset | base URL of a `grapharc serve --live-root` the requester can reach; posts a "watch live" link |

The timeout and live-interval environment values must be finite and positive;
fractions of a second are accepted. `NaN`, infinity and overflowing values
such as `1e309` are startup errors naming the variable. A requester-supplied
`--approval-timeout` must also be finite and positive and fit within the
command's existing timeout ceiling.

The bot reads tokens from the process environment only. The model gateway's
`.env` loader is deliberately not used here — even though it now reads the
working directory alone rather than searching upward: a bot that a whole
Expand Down
7 changes: 4 additions & 3 deletions docs/deep-dive.md
Original file line number Diff line number Diff line change
Expand Up @@ -139,7 +139,9 @@ Three things that phrase over-promises if left alone. **An interrupt does not st

**The HTTP API is FastAPI plus SSE** — create a session, list, get, post an event, stream the trace, fetch it as NDJSON, healthz. A request may name a registered graph and supply input and a budget; it may not *describe* a graph, because topology comes from a registry the operator fills in Python. But note the seam: **it does not use the session layer above.** It ships its own in-process runtime whose sessions die with the process, never evict, and record `message` and `approval` events without delivering them into a running graph. Two session layers that have not been joined ([ROADMAP.md](../ROADMAP.md) §12.3).

**Policy is a TOML document** over nodes, edges, tools and spend, with tiered evaluation — every `deny` before every `ask` before every `allow`, so a broad deny beats a narrow allow including one scoped to a single tenant. Every decision lands in an audit record naming the rule id, the reason, the policy version and a digest of the document, so a decision can be tied to the exact text that made it. And the seam, now narrowed to the tool plane: **the planner half is wired and the tool half is not.** `PolicyEngine.edge_policy()` and `PolicyEngine.node_policy()` compile the document into the `EdgePolicy` and `NodePolicy` the admission checker consults, and `grapharc plan --policy` is a real caller — so what may run, and what may connect to what, *is* governed by a document you can read. (A `resource = "node"` rule used to be dropped by the compiler and enforced by nothing; [issue #66](https://github.com/CodeGraphContext/GraphARC/issues/66).) But `permission_policy()`, `check_tool()` and `approval_router()` have no caller outside `grapharc/policy/`, so `grapharc agent` still assembles its tool gating from `--allow` / `--deny` / `--ask` globs. The most dangerous surface in the package is the one the document cannot reach yet; [issue #6](https://github.com/CodeGraphContext/GraphARC/issues/6) is that work, and the precedence question it has to settle is what happens when a flag `allow` meets a document `deny`.
**Policy is a TOML document** over nodes, edges, tools and spend, with tiered evaluation: every `deny` before every `ask` before every `allow`. The planner compiles node and edge rules into its admission gate. `grapharc agent --policy` compiles tool rules into the harness, combines them with CLI restrictions using `DENY > ASK > ALLOW`, and routes document approval requests through the policy's named role. An `--allow` flag cannot widen a document denial or bypass its approval requirement. JSON mode and redirected stdin refuse approval requests, and a selected policy document refuses delegated Claude CLI execution.

Agent policy and tenant settings follow `flag > GRAPHARC_* environment > grapharc.toml > default`, including `--config` and paths anchored to the config file. Governed results report the source, document version and digest, tenant, and policy audit path. Document-driven denials and approval decisions are audited; ordinary allowed tool calls and flag-only denials are not written to the document audit. The run trace still records attempted tool calls and their outcomes. These limits matter when assessing the remaining work in [issue #6](https://github.com/CodeGraphContext/GraphARC/issues/6).


## Independent verification
Expand Down Expand Up @@ -254,7 +256,7 @@ A stable system is not one that claims to have no edges — it is one whose edge
- **`.env` and `grapharc.toml` follow the same discovery rule: the working directory, and nowhere else.** Neither searches parent directories — a run must not be governed by a file you did not know about, and must not be *billed* to one either. **This is a behaviour change:** the credential loader used to walk up to `/`, so a `.env` in an ancestor directory (a `$HOME` one on a shared box, a client project one above a demo checkout) was picked up silently. If you relied on that, move the file into the directory you run from, `export` the variable, or pass `env_file=` to name it explicitly. A real environment variable still beats any file.
- **`grapharc run` has no budget unless you give it one.** Set any of `--max-tokens`, `--max-iterations`, `--max-seconds`, or `--max-concurrency`; without them each dimension is unlimited and the gate admits a topology of any worst-case cost.

**Verified this pass:** `pytest` → green, 2,252 selected and 13 deselected (the live ones); `ruff check .` clean; all eight `grapharc demo` stages green, plus the `trace` / `metrics` / `viz` / `replay` tour against a freshly recorded demo trace; the wheel builds and imports all submodules in a clean virtualenv with `[all]`, and `0.1.8` on PyPI is that wheel. The counts are a snapshot, not a property of the project — `pytest` re-derives them in one command, which is the only reason they are quoted, and `tests/test_deep_dive.py` fails this line rather than letting it drift.
**Verified this pass:** `pytest` → green, 2,305 selected and 13 deselected (the live ones); `ruff check .` clean; all eight `grapharc demo` stages green, plus the `trace` / `metrics` / `viz` / `replay` tour against a freshly recorded demo trace; the wheel builds and imports all submodules in a clean virtualenv with `[all]`, and `0.1.8` on PyPI is that wheel. The counts are a snapshot, not a property of the project — `pytest` re-derives them in one command, which is the only reason they are quoted, and `tests/test_deep_dive.py` fails this line rather than letting it drift.

[ROADMAP.md](../ROADMAP.md) tracks what is built and what is not, item by item.

Expand All @@ -265,4 +267,3 @@ Defects that have been **closed** — each with what broke, how it was found and

Architecturally *inspired by* systems studied from public documentation: OpenClaw (policy-before-schema tool gating, file-first state, and its security post-mortems), Hermes Agent (budgeted tiered memory, ephemeral subagents), Claude Code (advisory-vs-enforced split, subagent context isolation, verification-centered loops), and OpenRouter (routing semantics, budget-scoped accounting).


Loading
Loading