diff --git a/README.md b/README.md index 04aff22..18ded5e 100644 --- a/README.md +++ b/README.md @@ -10,7 +10,7 @@ The engineering principles and playbooks remain pstack's. The runtime integratio - Codex browser, computer-use, GitHub, and automation tools - `.codex/skills` for global and project-local skills -The port tracks upstream **0.15.10** at [`4e5b1cf`](https://github.com/cursor/plugins/commit/4e5b1cf2ccb0ea3716f08c8ee0a5856b5ab93536), reviewed on October 5, 2026. It contains **50 skills**, including 24 engineering principles. This update adds `poteto-help`, `correct`, `benchmark-checklist`, three principle skills, a checked multi-PR plan, stronger architecture screening, and evidence-backed benchmark and PR verification guidance. See [the port record](docs/upstream-port.md) for adaptations and exclusions. +The port tracks upstream **0.15.15** at [`ccb5507`](https://github.com/cursor/plugins/commit/ccb5507cec1546dc88135c1139c811e6c59115ba), reviewed on October 9, 2026. It contains **50 skills**, including 24 engineering principles. This update adds prompting references, copyable recipes, missing-configuration guidance, and explicit reasoning-budget targets. The Codex-only Sol, Luna, and Astra routes remain the balanced defaults. See [the port record](docs/upstream-port.md) for adaptations and exclusions. ## Install @@ -69,6 +69,10 @@ gh auth status The watcher installs its pinned dependencies on first use and reads GitHub state without granting merge authority. Babysit stops at merge-ready. Shipping requires an explicit request to merge, land, or enable merge when ready. +Read the [Codex usage guide](docs/usage.md) for prompting, verification, longer runs, and repeated mistakes. Ask `Use poteto-help` for one prompt tailored to your task. Help answers the question without starting that workflow. + +Budget profiles are explicit targets: balanced keeps the configured role efforts, unlimited targets max, large targets xhigh, medium targets high, and small targets medium. Setup preserves model choices and uses the highest supported effort at or below the target. Inherited aliases stay inherited. Updating the skills preserves your existing global model configuration; rerun setup only when you want to change it. + ## Check upstream ```bash diff --git a/UPSTREAM_COMMIT b/UPSTREAM_COMMIT index 0180894..8951d5b 100644 --- a/UPSTREAM_COMMIT +++ b/UPSTREAM_COMMIT @@ -1 +1 @@ -4e5b1cf2ccb0ea3716f08c8ee0a5856b5ab93536 +ccb5507cec1546dc88135c1139c811e6c59115ba diff --git a/config.example.md b/config.example.md index 060e4c5..37e6101 100644 --- a/config.example.md +++ b/config.example.md @@ -9,8 +9,8 @@ ## Budget - profile: balanced -- policy: preserve the per-role efforts below; a requested large, medium, or small budget maps routes to xhigh, high, or medium only when that model supports it -- verified: 2026-10-05 from the active Codex collaboration schema +- policy: balanced preserves the per-role efforts below; unlimited, large, medium, and small target max, xhigh, high, and medium respectively, using a supported effort at or below the target +- verified: 2026-10-09 from the active Codex collaboration schema ## Model routes diff --git a/docs/upstream-port.md b/docs/upstream-port.md index c41eace..0b5a944 100644 --- a/docs/upstream-port.md +++ b/docs/upstream-port.md @@ -1,5 +1,23 @@ # Upstream port record +## October 9, 2026: upstream 0.15.15 + +Reviewed all five pstack commits from `4e5b1cf2ccb0ea3716f08c8ee0a5856b5ab93536` through [`ccb5507cec1546dc88135c1139c811e6c59115ba`](https://github.com/cursor/plugins/commit/ccb5507cec1546dc88135c1139c811e6c59115ba). The latter is the reviewed repository head; the latest pstack-specific commit is `df58112`. + +| Upstream commit | Disposition | +|---|---| +| [00b52d9](https://github.com/cursor/plugins/commit/00b52d9) | Keep the help/work boundary. Do not import Cursor's typed-only frontmatter; this port preserves Codex's existing discovery policy and name/description frontmatter. | +| [807c031](https://github.com/cursor/plugins/commit/807c031) | Port prompting and recipe references. Adapt recurring work to requested Codex heartbeat automations, constrain worker fan-out, and preserve authorization boundaries while away. | +| [2cbf585](https://github.com/cursor/plugins/commit/2cbf585) | Adapt the guide additions into `docs/usage.md` and the help references: evidence, prototypes, verifiable plans, repeatable harnesses, benchmarks, repeated mistakes, and longer-run handoffs. Exclude Cursor cloud projects, Bot UI, and the upstream Slack automation pack. | +| [1e56b29](https://github.com/cursor/plugins/commit/1e56b29) | Offer setup now or later once per chat when configuration is missing and relevant, and review seeded defaults during onboarding; answer the original question either way and preserve existing choices. | +| [df58112](https://github.com/cursor/plugins/commit/df58112) | Port explicit budget targets: unlimited now means max, with supported-effort fallback. Preserve balanced Codex-only model routes and three-seat panels. Exclude the switch to non-Codex providers, related wording churn, Cursor plugin metadata, and the excluded orchestration test fixture rename. | + +The Codex-only routes still match the active collaboration schema. Balanced keeps the existing efforts. Updating skills does not migrate a user's persisted configuration or silently raise their budget. The help references and usage guide explain actual Codex behavior rather than claiming Cursor invocation controls or cloud isolation. + +Validation for this update includes the repository audit, existing unit tests, skill validation for the changed skills, and link checks for the new documentation. Global installation is verified against repository content and preserves existing model configuration. These are packaging checks, not claims of end-to-end workflow evaluation. + +## October 5, 2026: upstream 0.15.10 + This update ports the pstack changes between `9490cc1cf95d5de2e4941196cdac00dd861812a4` and [`4e5b1cf2ccb0ea3716f08c8ee0a5856b5ab93536`](https://github.com/cursor/plugins/commit/4e5b1cf2ccb0ea3716f08c8ee0a5856b5ab93536). The reviewed upstream version is 0.15.10, checked on October 5, 2026. All 25 pstack commits in that range were reviewed. The original MIT license and attribution remain. Codex runtime instructions take precedence over upstream Cursor integration. diff --git a/docs/usage.md b/docs/usage.md new file mode 100644 index 0000000..a129539 --- /dev/null +++ b/docs/usage.md @@ -0,0 +1,59 @@ +# Using pstack in Codex + +Start a new Codex chat in the repository you want to work on. For substantial work, say `Use poteto-mode` followed by the goal and a check that can pass or fail. For an obvious small edit, ask directly. Installation makes skills available; it does not start workflows or schedule background work. + +## Give the agent a checkable task + +A useful prompt gives the goal, done check, proof to return, known symptoms, and real constraints. For example: + +```text +Use poteto-mode. The CSV export drops its last row since yesterday's deploy. +The failing job is 4812. Reproduce first, then fix it. Done means the +60,000-row fixture exports every row. Show row counts before and after. +``` + +For a noisy report, first ask the agent to restate the problem in plain words. Share observations and logs, and identify theories as hypotheses. Once the chat has enough context, "continue" is enough. Say "new task" when changing subjects. + +The [prompting reference](../skills/poteto-help/references/prompting.md) explains how to steer a run. The [recipes](../skills/poteto-help/references/recipes.md) cover investigation, design, implementation, review, and handoff. `Use poteto-help` returns guidance and a prompt you can send; it does not start the example task. + +## Understand and design before changing code + +When the cause is unclear, ask for a read-only investigation with findings, evidence, and hypotheses. `how` traces current behavior; `why` investigates the reasons behind it. `recall` can recover scoped prior context from available history. `teach` explains the mechanism and tradeoffs in plain language. + +For consequential design choices, ask for a few runnable prototypes. Use screenshots for UI, or real outputs and measurements for behavior. For a shared package or API, write the caller's tutorial first and work back to the implementation. `architect` handles boundaries and types; `arena` compares alternatives. `swarm` covers independent slices with bounded workers. `interrogate` challenges a diff without applying its findings automatically. + +Once the design is settled, ask poteto-mode to turn it into small verifiable PRs through the Multi-phase plan playbook. Planning returns a plan; execution needs a request to proceed. For a compatibility migration, state whether output must match exactly, including existing bugs. + +## Make verification repeatable + +Match proof to the changed behavior: run the real CLI, drive the UI flow, replay an input, compare profiles, or read back persisted state. Ask for an artifact you can inspect. A green build alone does not prove the user-visible result. If a fix is already merged, repeat the relevant check on main. + +Use an existing project verification harness when available. If none exists, `create-verification-skill` can create a project-local harness when installed and requested. Verify it end to end before relying on it. Keep its feature map current with `maintain-verification-skill` when available; request a scheduled review if that cadence helps your project. + +Repeated manual app-driving scripts can become a small control CLI with useful help, machine-readable output, actionable errors, and dry runs for destructive operations. Seeded data and one repeatable development command make verification easier for every agent. + +For performance claims, run `benchmark-checklist`. It checks the limiting resource, production tuning, physical limits, correctness, alternating measurements, end-to-end impact, and whether the timed work actually ran. An unexplained or untuned comparison is inconclusive. Runtime forensics diagnoses a live process; Trace forensics maps an existing profile to source. Ask for diagnosis only when a fix is not yet wanted. + +## Give longer runs clear scope + +Before leaving a task unattended, observe a successful run, give the agent the verification tools it needs, and make each stage produce evidence. State a checkable done condition, allowed actions, review gates, and what to do if blocked. Request a decision log for later review. + +Give concurrent writers separate worktrees or output paths. Codex collaboration workers share a machine, so worktrees alone do not isolate ports or build caches. Respect the configured concurrency cap and give conflicting processes separate resources. + +For later or recurring work, explicitly request a Codex heartbeat automation when available. It should stay quiet while nothing meaningful changes and report completion, failure, or required input. Confirm that it was actually saved. A local run does not imply that work continues while the machine is unavailable. + +Say "pause safely" to request a checkpoint and resume note. On return, ask poteto-mode to take over the branch and read that note. "Keep going" continues the current authorized work. A plan, an overnight prompt, or installing pstack does not itself authorize external messages or deployment. + +PR status, merge-ready, and merged are different outcomes. "Check on PR 123" makes one status pass. "Babysit PR 123 until merge-ready" follows checks and review feedback. "Merge PR 123" grants that merge scope. + +## Fix repeated mistakes in the repository + +`correct` finds repeated mistake classes in commits and review history. It prefers architecture and data structures that prevent the mistake, then types or checks, then tests, and finally written rules. Each new check should reject a real past mistake. `reflect` instead proposes skill improvements from a session; proposals need authorization before application. + +Add skills or checks in response to demonstrated failures. For a substantial skill change, ask for an evaluation against realistic tasks and inspect the outputs yourself. A structural skill validator checks packaging, not future agent judgment. + +## Control the models and cost + +The parent model is selected in Codex. The [configuration](../config.example.md) supplies compatible child routes and concurrency limits. Existing global choices survive installation. Smaller reasoning budgets and shorter panels can reduce work; inheritance uses the parent's model and is not inherently cheaper. + +`setup-pstack` can change models and reasoning budgets when requested. Balanced preserves per-role efforts; unlimited, large, medium, and small target max, xhigh, high, and medium. Unsupported efforts fall back within the same verified model at or below the target. Missing configuration or first-time onboarding is an opportunity to offer setup, not a reason to block the original question. The installer seeds defaults, so having a file does not prove you have chosen its settings. diff --git a/skills/poteto-help/SKILL.md b/skills/poteto-help/SKILL.md index 32550ac..f006f74 100644 --- a/skills/poteto-help/SKILL.md +++ b/skills/poteto-help/SKILL.md @@ -13,6 +13,8 @@ Install from [pstack-codex](https://github.com/HustleCoding/pstack-codex) with ` Run **setup-pstack** to choose verified Codex models, reasoning efforts, and bounded fan-out. It writes `~/.codex/pstack/config.md`. Existing settings remain intentional until the user asks to change them. A missing file means children inherit the session unless a workflow supplies a verified route. +If the configuration is missing and the question concerns setup, cost, or which models will run, offer setup now or later at most once per chat. For onboarding, offer to review the current model and budget choices even when the file exists: a fresh installer seeds defaults, so file existence does not prove the user chose those settings. The once-per-chat limit covers both offers. Answer the original question either way. If the user has already asked to configure it, proceed within that request. If they defer, explain inheritance and any verified workflow defaults. Do not replace an existing configuration merely because upstream defaults changed. + Suggested prompt. "Use setup-pstack. Configure balanced Codex-only routes and explain the model choices." Use **poteto-mode** with a concrete goal and a check that can pass or fail. Invoke it at the start of each new task. For a persistent project preference, the user can ask to add a scoped instruction to AGENTS.md. Do not assume a Custom Mode is available. @@ -21,6 +23,10 @@ Suggested prompt. "Use poteto-mode. Diagnose this bug, prove the cause, fix it, The parent model is chosen in Codex's model picker. A route cannot change the active parent's model. Child overrides require a compatible collaboration schema and a fresh child with minimal task-local context; full-history forks inherit. When overrides are unavailable, report inheritance. +## Help with a prompt + +Read [references/prompting.md](references/prompting.md) before helping word a task or steer a drifting run. It covers goals, done checks, evidence, context, and constraints. Use [references/recipes.md](references/recipes.md) for a matching copyable example. Lead with the answer, give at most one adapted prompt unless the user asks for more, and link the skill or reference that supports it. Help does not start the example workflow. + ## Pick a skill | Goal | Skill | diff --git a/skills/poteto-help/references/prompting.md b/skills/poteto-help/references/prompting.md new file mode 100644 index 0000000..cf09589 --- /dev/null +++ b/skills/poteto-help/references/prompting.md @@ -0,0 +1,51 @@ +# Word the prompt + +A prompt states the intent and the check for done. The playbook supplies the steps, so a few plain sentences beat a spec. + +## Put in + +- The goal. Say what is wrong or what the user wants. +- The done check. It can pass or fail. "Make it better" and a duration are not checks. +- The proof to show. Ask for the real command output, a video of the flow, the stored value, or a before and after number. +- What the user already knows. A symptom, a repro step, a log, or a link saves the agent a search. +- The real constraints. "repro first", "don't change any code yet", "zero behavior change", and "let me review before proceeding" each change what the agent does. + +## Leave out + +- The how. Say what to achieve, and leave the agent room to find a better way. +- A list of skills or steps. A hand-written order drops or reorders steps the playbook keeps. Name a skill only to override one choice. +- The user's theory of the cause, until the agent restates the problem. A stated guess narrows the search. + +## Load the context first + +- For a noisy report, ask the agent to restate the underlying issue in its own words and in plain English before it does anything else. A misreading shows up before any code exists. +- In a fresh chat, `/recall` earlier work on the topic. Old chats hold context that the new agent lacks. +- Before a change to unfamiliar code, ask `/how` for the mechanics and `/why` for the reasons. An agent with no traced model fixes the symptom at the first plausible spot. +- Ask `/teach` to make the case for a choice, as in "convince me it fixes the cause and not the symptom". A case is easier to check than a summary. + +## Design before the plan + +- For a consequential design choice, ask for prototypes of a few options, with screenshots or videos for UI, and pick from the evidence. +- Let prototypes answer the open questions. Use runnable evidence to settle questions before asking reviewers to challenge the implementation. +- For a shared package or API, ask for the README or a tutorial first, then work back to the code. The doc becomes the target the agent checks itself against. +- Ask for the plan only after the design is settled. Each step of the plan ends in a check. + +## Follow up short + +- "do it", "continue", and "keep going until done" are whole prompts once the chat holds the task. +- Start with "new task" when the subject changes. Otherwise the mode treats the message as the next step. + +## Before stepping away + +- State the decisions the agent may make while you are away and any review gates it must preserve. Stepping away does not expand authorization. +- Write done as checks every iteration can run, and request a Codex heartbeat automation with that predicate when later or recurring work is needed. +- Ask for a fresh worktree off a named base. +- Pre-answer what the agent would stop for, such as "don't ask me before committing". +- Ask for a decision log to audit later. +- Give an exit: "if you're truly stuck after a few hours, stop and write up why". + +## Steer in one line + +- Restate the goal: "i said the goal is to repro. i did not ask for a fix yet." +- Name the principle: "apply prove it works. show me the real output, not the build log." +- A principle name works because the agent already read the rule. Its reply names the decision the rule changed. diff --git a/skills/poteto-help/references/recipes.md b/skills/poteto-help/references/recipes.md new file mode 100644 index 0000000..6f099b9 --- /dev/null +++ b/skills/poteto-help/references/recipes.md @@ -0,0 +1,49 @@ +# Prompts worth copying + +Swap in the real paths, skills, and done checks. Informal wording works. In Codex, invoke a skill with `Use ` or an explicit skill mention. The slash names below identify skills; playbooks are selected through poteto-mode. Verification skills are project-specific: use one only when installed. + +A prompt grants only its stated scope. Publishing, merging, and scheduled follow-up are separate actions; the examples name them where intended. + +## Understand + +- `/poteto-mode read . restate the underlying issue in your own words, in plain english.` +- `/poteto-mode investigate why . give me what we know, what data you used, and your best hypotheses. don't change any code yet.` +- `use /how to understand . then use /why to find out why it broke recently.` +- `/recall my work on from last week, then read .` +- `/teach me why you implemented it this way and not . what did you trade off?` +- `/poteto-mode take over this branch. read the decision log, find what's done, and continue. don't redo finished work.` + +## Build + +- Bug: `/poteto-mode . repro first, then fix and verify.` +- Bug in an app: `/poteto-mode repro this with /verify-. if it repros on main, fix it and show me a video as proof.` +- Bug with a cheap test: `/poteto-mode repro first. if there's a cheap test path, /tdd it. then fix and rerun.` +- Feature: `/poteto-mode add . stays byte-identical. verify both.` +- Refactor: `/poteto-mode move into one module, zero behavior change. record the current output first and prove it's unchanged after.` +- Perf: `/poteto-mode takes