Repository navigation
feat: add scoped Spec Studio audits and confirmed remediation - #260
Coding-Dev-Tools wants to merge 9 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3a7e1dd391
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (node_a["subject_key"] and node_b["subject_key"] | ||
| and node_a["subject_key"] != node_b["subject_key"]): | ||
| continue | ||
| if (node_a["claim_kind"] and node_b["claim_kind"] | ||
| and node_a["claim_kind"] != node_b["claim_kind"]): |
There was a problem hiding this comment.
Require a shared subject before reporting parameter conflicts
When either record lacks a subject_key or claim_kind, these guards allow any two memories with the same parameter name and different values to become a high-severity conflict. For example, an unrelated web-server port: 8000 and database port: 5432 are presented as contradictory with a one-click supersession remedy, even though no common subject was established. Require positively matching claim identity or other subject evidence before proposing retirement.
Useful? React with 👍 / 👎.
| newer = node_a if node_a["ingested_at"] >= node_b["ingested_at"] else node_b | ||
| older = node_b if node_a["ingested_at"] >= node_b["ingested_at"] else node_a |
There was a problem hiding this comment.
Choose conflict keepers by valid time
For backfilled or imported facts, ingestion order does not represent world-time recency: a historical record can be ingested after the currently valid record. Selecting the keeper solely by ingested_at can therefore recommend retaining the historical value and retiring the current one; use valid_from (with a deliberate fallback) or require the user to choose when temporal ordering is ambiguous.
AGENTS.md reference: AGENTS.md:L175-L177
Useful? React with 👍 / 👎.
| except Exception: | ||
| pass | ||
| return None |
There was a problem hiding this comment.
Preserve unchecked status when claim lookup fails
If grounded_recall raises because the vector index or store is temporarily unavailable, this wrapper swallows the exception and returns None; crawl_spec interprets that as a completed lookup with no evidence, marks the claim untraced, and penalizes the score. Let the exception reach crawl_spec (which already converts lookup failures to unchecked) or return a distinct failure result so infrastructure errors are not reported as unsupported claims.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 99189b0990
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if act == "supersede": | ||
| self.retire(node_b_clean, workspace=ws or "default", repo=route["repo"], | ||
| reason="spec_crawl_resolve", actor="spec_crawl") |
There was a problem hiding this comment.
Close the retired fact at the replacement's effective time
For a backfilled conflict where the retained node's valid_from predates the remediation, this calls ordinary retire(), which closes the old node at the current wall-clock time rather than at the replacement's effective time. Although the remedy now selects the correct keeper by valid_from, an as_of read between that effective time and the remediation time still returns both contradictory facts; close validity at the retained node's valid_from instead.
AGENTS.md reference: AGENTS.md:L175-L177
Useful? React with 👍 / 👎.
| for mid in (node_a_clean, node_b_clean): | ||
| self._check_owns(mid, wid, rid) |
There was a problem hiding this comment.
Permit ancestor nodes in repo-scoped remediation
When a repo-scoped audit includes a workspace/user ancestor, as its SearchFilter(include_ancestors=True) intentionally does, _check_owns(..., rid) rejects that ancestor because its repo_id is None. Consequently every generated supersede or link remediation involving a repo node and an ancestor fails with “does not belong to that workspace/repo”; validate both nodes against the same ancestor-aware audit scope and avoid reapplying exact-repo ownership during the mutation.
AGENTS.md reference: AGENTS.md:L178-L179
Useful? React with 👍 / 👎.
| for mid in mem_ids_set: | ||
| node_links = self.store.get_links(mid) | ||
| for nl in node_links: | ||
| if nl.get("a") in mem_ids_set and nl.get("b") in mem_ids_set: | ||
| links.append(nl) |
There was a problem hiding this comment.
Deduplicate stored links before cluster analysis
For every stored A–B link, both get_links(A) and get_links(B) return the same row, so this loop appends every edge twice. Besides duplicating the API's links and replay events, these duplicates inflate resolved_links, whose length is used to decide whether semantic-link inference runs; for example, two stored edges among four nodes appear as four links and suppress inference entirely. Deduplicate by endpoints and relation before calling analyze_memory_nodes.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d87bb31d9c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| elif _XML_CLOSE.match(line): | ||
| title = "" # closing tag: consumes the line, starts no section | ||
| if title is not None: | ||
| heads.append((offset, min(line_end + 1, len(text)), title)) |
There was a problem hiding this comment.
Preserve text after XML closing tags
When a spec uses XML-style sections, adding the closing tag as an empty heading ends the preceding section but creates no section for subsequent unheaded text. For example, content after </role> is never tokenized, so trailing rules, vague wording, or injection payloads are omitted and the report can incorrectly appear clean. Preserve that trailing span as a section instead of silently dropping it.
Useful? React with 👍 / 👎.
| if target not in _ENFORCE_VERBS | _FORBID_VERBS | _STOPWORDS: | ||
| directives[target] = action |
There was a problem hiding this comment.
Keep outer negation from being overwritten
For directives such as Never require authentication or Never enforce deployment, the outer never first records a forbid directive, but the scan later processes require/enforce independently and overwrites the same target with enforce. As a result, comparison with Always require authentication reports no conflict, hiding a direct policy contradiction; preserve the governing negation when consuming nested directive verbs.
Useful? React with 👍 / 👎.
| "title": "Crawl spec or prompt quality", | ||
| "readOnlyHint": True, | ||
| "destructiveHint": False, |
There was a problem hiding this comment.
Register crawl-only MCP tools as viewer reads
In hosted/team MCP usage, minimum_role() does not derive access from this annotation and defaults unlisted tools to member. Because neither engraphis_spec_crawl nor engraphis_spec_crawl_memories was added to _READ_ONLY_TOOLS, viewers are denied these explicitly read-only audit operations even though they can use the existing viewer-level memory reads; register both tools in the read-only role set.
Useful? React with 👍 / 👎.
Description
Spec Studio adds a deterministic v2 specification crawler and scoped memory-cluster audit to the dashboard, REST API, Classic MCP, and Smart MCP discovery. It reports ambiguity, injection patterns, claim support, and advisory memory remediation while leaving analysis read-only.
Parameter conflicts require positively matching nonempty subject keys and compatible claim kinds. The index groups records by subject before comparison so unrelated parameter values cannot create or hide conflicts. Supersession suggestions use distinct finite
valid_fromdates; ingestion order cannot select a historical backfill as the keeper. Missing, invalid, or equal effective dates and unproven divergences produce clarification without keeper/retirement targets. The UI displays those instructions and effective dates safely. Failed evidence lookups remainunchecked, do not count as unsupported claims, and do not penalize the score; successful abstentions remainuntraced.Scope and owner checks restrict memory analysis to approved live records. HTTP accepts direct text or approved procedural-memory inputs; local-operator MCP file reads stay within approved roots. Clarifications remain pending for review. Supersession requires explicit confirmation and validates both records before atomically linking and closing validity. The UI rejects stale scoped responses, renders untrusted text safely, reports actual failures, and supports accessible clarification dialogs.
The full browser gate exposed a deferred-fit camera race in the every-node renderer. Hidden zero-size measurements now preserve the last usable viewport. Worker preview, ready, and final-layout fitting respect manual keyboard, wheel, drag, pinch, and focus navigation; explicit Fit remains available. Controlled real-worker browser tests preserve exact projected camera equality across hiding and replay, including tiny drags and ordinary automatic fitting.
Evaluation enforces section counts and score bounds in full-stack and NumPy-only CI. MCP contracts, skill hashes, guides, and public charts are synchronized with immutable offline evidence. README benchmark links are pinned to the source/evidence commit for PyPI rendering. Historical artifacts and unrelated primary-checkout work remain preserved.
Type
Verification
ruff check ., Pyright 1.1.414, commercial/CSP drift checks, MCP contract export, and the Spec Crawl evaluation passed.Current head:
d87bb31d9c09265904b2c4477dda9831c7ecf5db. Fresh full CI is running at https://github.com/Coding-Dev-Tools/engraphis/actions/runs/38070240255; results from the previous head do not qualify this changed head. The three review findings are implemented and their original threads are outdated; formal thread resolution remains subject to the repository approval protocol.