Compare commits

...

9 Commits

Author SHA1 Message Date
723b7be84a change: send down proj context with pi tool/skills 2026-08-05 10:58:05 -04:00
bed74f9d33 Make /ralpi-reset select the loop to reset like resume 2026-08-03 20:18:54 -04:00
3a3c293aa7 bump 2026-08-03 17:13:52 -04:00
b9a12931fe fix: resume selection under-reports task totals with multiple loops
The resume prompt counted total/completed from progress.tasks, which only
records TOUCHED tasks (started/completed/failed) — never-started tasks were
silently missing, and file-checkbox completions weren't counted unless
markCompleted had run. With e.g. 10 tasks where 3 completed and 5 untouched,
it showed 3/5 done instead of 3/10 done.

- add countPRDResumeStats (src/utils.ts): total from parseTaskFile of the
  PRD source; completed = progress completions ∪ PRD checkbox completions,
  deduped by task id; failed from progress. Falls back to touched-task
  counts when the source file is missing/unparseable.
- selectPRDToResume (index.ts) uses the helper instead of Object.keys().
- tests/resume-stats.test.ts: 4 regression tests.

Includes pre-existing uncommitted changes: index.ts tab-reformat (matches
src/ style) and package.json version bump 0.4.2 -> 0.4.3.
2026-08-03 16:31:09 -04:00
f9b57ec2ed fix flaky resume behavior 2026-08-02 23:36:35 -04:00
b0749467c9 fix: updated docs to align with reality 2026-08-02 16:54:27 -04:00
f361f05f96 version bump 2026-08-02 16:43:42 -04:00
9f1250d1a1 fix: ralpi-plan properly loads task-manager prompt 2026-08-02 16:43:08 -04:00
fc2f8a879f bump 2026-07-31 09:58:22 -04:00
13 changed files with 2773 additions and 1627 deletions

100
AGENTS.md
View File

@@ -2,7 +2,9 @@
## What this is ## What this is
A Pi coding agent extension that registers the `/ralpi` slash command. Not a standalone app — it runs inside Pi's extension host. A Pi coding agent extension that registers the `/ralpi` slash commands
(`/ralpi`, `/ralpi-run`, `/ralpi-plan`, `/ralpi-resume`, `/ralpi-reset`).
Not a standalone app — it runs inside Pi's extension host.
## Type checking ## Type checking
@@ -10,6 +12,8 @@ A Pi coding agent extension that registers the `/ralpi` slash command. Not a sta
npm run typecheck # tsc --noEmit npm run typecheck # tsc --noEmit
``` ```
Tests: `bun test` (parser and DAG suites in `tests/`).
No build step needed — Pi loads extensions via [jiti](https://github.com/unjs/jiti), which compiles TypeScript at runtime. `index.ts` is the entry point directly. No build step needed — Pi loads extensions via [jiti](https://github.com/unjs/jiti), which compiles TypeScript at runtime. `index.ts` is the entry point directly.
## Entry point ## Entry point
@@ -23,22 +27,41 @@ The extension imports from Pi SDK packages (not in `package.json` — provided b
- `@earendil-works/pi-coding-agent``ExtensionAPI`, `ExtensionContext`, `createAgentSession`, etc. - `@earendil-works/pi-coding-agent``ExtensionAPI`, `ExtensionContext`, `createAgentSession`, etc.
- `@earendil-works/pi-tui``Box`, `Text` for custom message renderer - `@earendil-works/pi-tui``Box`, `Text` for custom message renderer
The only real npm dependency is `yaml` (^2.4.0). The only real npm dependency is `yaml` (^2.4.0). It is used for parsing YAML
task files (`src/parser.ts`) and config files (`parseSimpleYaml` in
`src/utils.ts`, which falls back to a flat key:value parser when the package
is unavailable).
## Source structure ## Source structure
- `index.ts` — extension entry, command routing, UI registration, reload detection - `index.ts` — extension entry, command registration (`ralpi`, `ralpi-run`,
`ralpi-plan`, `ralpi-resume`, `ralpi-reset`), execution-mode + loop-option
prompts, reload auto-resume via `session_start`, progress message renderer
- `src/` — all logic modules: - `src/` — all logic modules:
- `parser.ts` — task file parsing (Fio, checkbox, YAML formats) - `parser.ts` — task file parsing (Fio/README numbered, phased, checkbox,
- `dag.ts` — Kahn's algorithm dependency resolution, batch planning YAML formats), dependency + parallel-group + timeout parsing,
- `executor.ts` — task execution, retry, parallel/sequential modes `updateTaskInFile()` for PRD checkbox updates
- `progress.ts``.ralpi/progress.json` state management - `dag.ts` — Kahn's algorithm dependency resolution, group-aware batching,
cycle detection, sequential/parallel plan builders
- `executor.ts` — task execution, parallel/sequential modes, model
round-robin + failover, review-gated loop, worktree orchestration,
batch-level conflict resolution
- `review.ts` — review verdict extraction (`## REVIEW VERDICT`), review
save/load to `.ralpi/reviews/`
- `worktree.ts` — git worktree create/merge/cleanup helpers, stale-worktree
cleanup, `finalizeCommittedWorktrees()`
- `progress.ts``.ralpi/progress.json` state management (multi-PRD)
- `prompts.ts` — prompt generation for spawned agent sessions - `prompts.ts` — prompt generation for spawned agent sessions
- `reflection.ts` — reflection extraction from agent output - `reflection.ts` — reflection extraction from agent output
- `utils.ts` — config loading, progress discovery, `runAgentSession()` - `utils.ts` — config loading, progress/PRD discovery, `runAgentSession()`,
model resolution (`resolveModelSpec`), loop-active marker
- `types.ts` — all interfaces and `DEFAULT_CONFIG` - `types.ts` — all interfaces and `DEFAULT_CONFIG`
- `widget-batcher.ts` — debounced widget updates for parallel tasks - `widget-batcher.ts` — debounced widget updates for parallel tasks
- `constants.ts` — static constants - `task-manager-prompt.ts` — loads and expands the bundled
`prompts/task-manager.md` template for `/ralpi-plan`
- `constants.ts` — static constants (slash command, task file names,
reflection/review patterns)
- `tests/` — bun test suites for parser and DAG behavior
- `skills/ralpi-use.md` — Pi skill definition for task execution - `skills/ralpi-use.md` — Pi skill definition for task execution
- `prompts/task-manager.md` — Pi prompt for task planning - `prompts/task-manager.md` — Pi prompt for task planning
@@ -47,27 +70,70 @@ The only real npm dependency is `yaml` (^2.4.0).
All runtime state lives in `.ralpi/` in the **project directory** (not this extension directory): All runtime state lives in `.ralpi/` in the **project directory** (not this extension directory):
- `.ralpi/progress.json` — execution progress, supports multiple PRDs - `.ralpi/progress.json` — execution progress, supports multiple PRDs
- `.ralpi/loop-active.json` — marker written while a loop runs; drives
auto-resume after a session reload
- `.ralpi/reflections/` — per-task reflection JSON files - `.ralpi/reflections/` — per-task reflection JSON files
- `.ralpi/reviews/<prdKey>/` — full review output JSON (only when
`saveReviews` is on)
- `.ralpi/prompts/` — generated prompts (timestamped, for debugging) - `.ralpi/prompts/` — generated prompts (timestamped, for debugging)
- `.ralpi/sessions/` — full session transcripts - `.ralpi/config.yaml` — project-level config (optional)
There is no `.ralpi/sessions/` directory anymore — full task output is shown
inline via expandable `ralpi-progress` chat messages, and review output is
persisted under `.ralpi/reviews/`.
## Task ID convention ## Task ID convention
Task IDs are zero-padded strings (`"01"`, `"02"`, etc.). The parser prepends `0` to parsed digits. Never use raw numeric IDs. Task IDs are zero-padded strings (`"01"`, `"02"`, etc.) with an optional
single lowercase letter suffix for sub-tasks (`"02b"`, `"02c"`). The parser
normalizes `2b``02b` (see `normalizeTaskId` in `src/parser.ts`). Never
use raw numeric IDs.
## Command routing ## Command routing
`/ralpi` with no args → plan. First token looks like a path (`@path`, `./path`, `.md`, etc.) → run. Otherwise dispatches to subcommand (`run`, `plan`, `resume`, `reset`). - `/ralpi` no args → show plan for `README.md`; first token looks like a
path (`@path`, `./path`, `.md`, `.yaml`, etc.) → run; anything else →
error suggesting the dash commands
- `/ralpi-run [task-file]` — run tasks (auto-resumes when progress already
exists for the file; otherwise prompts for execution mode + loop options)
- `/ralpi-plan [prompt]` — loads the bundled `prompts/task-manager.md`
template and sends it as a user message. Pi's `sendUserMessage()` sends
with `expandPromptTemplates: false`, so the extension does its own
frontmatter stripping and `$@`/`$1` arg substitution
(`loadTaskManagerPrompt` in `src/task-manager-prompt.ts`)
- `/ralpi-resume [task-file]` — resume from persisted progress; prompts for
the PRD when multiple loops have progress. Reuses the loop snapshot in
`loop-active.json` (mode + autoCommit/autoReview/saveReviews) to resume
non-interactively
- `/ralpi-reset [task-file]` — reset execution progress (does not modify the PRD)
The old `/ralpi plan|resume|reset` subcommand dispatch, plus `status` and
`next`, were removed.
## Config ## Config
Read from `.ralpi/config.yaml` in project directory (and global `~/.pi/ralpi/config.yaml`). Falls back to `DEFAULT_CONFIG` in `src/types.ts` when files are missing. Config is loaded at `projectDir` level, not extension level. Read from `.ralpi/config.yaml` in project directory (and global
`~/.pi/ralpi/config.yaml`), project overrides global. Falls back to
`DEFAULT_CONFIG` in `src/types.ts` when files are missing. Config is loaded
at `projectDir` level, not extension level. Execution keys explicitly
present in a loaded YAML are tracked in `execution.explicitKeys` so the
loop-startup interactive prompts (`selectLoopOptions` in `index.ts`) can be
skipped for fields the user already set.
Key config fields in `execution`: Key config fields in `execution`:
- `autoCommit` / `autoReview` — toggle follow-up commit and review agent sessions (also selectable at loop startup via `selectLoopOptions`) - `autoCommit` / `autoReview` / `saveReviews` — loop options (selectable at
- `models` — round-robin model list for parallel mode loop startup via `selectLoopOptions`; review is asked FIRST, commit is
- `implModel` / `commitModel` / `reviewModel``<provider>/<model>` strings resolved via `resolveModelSpec` in `utils.ts` mandated when review is on)
- `models` — slot-aware round-robin model list for parallel mode, with
automatic failover to the next model per task
- `implModel` / `commitModel` / `reviewModel``<provider>/<model>` strings
resolved via `resolveModelSpec` in `utils.ts`
- `maxReviewRetries` / `reviewBlockOnFail` — review-gated loop retry behavior
- `worktrees``"never" | "parallel" | "always"` git worktree isolation
(default `"parallel"`; see `shouldUseWorktrees` in `src/executor.ts`)
- `commitTimeoutMs` / `reviewTimeoutMs` — timeouts for follow-up sessions - `commitTimeoutMs` / `reviewTimeoutMs` — timeouts for follow-up sessions
- `loopTimeoutMs` — max total loop duration in ms (0 = no limit; checked between batches in `executePlanBatches`) - `loopTimeoutMs` — max total loop duration in ms (0 = no limit; checked
between batches in `executePlanBatches`)
- `timeoutMs` — per-task execution timeout - `timeoutMs` — per-task execution timeout
- `prompts.projectContext` / `prompts.reflectionPrompt` — prompt-level settings

189
README.md
View File

@@ -8,35 +8,59 @@ pi install npm:@mikefreno/ralpi
## Features ## Features
- **Parallel batching**: Independent tasks in each batch can run concurrently - **DAG-based execution**: Tasks ordered via dependencies (arrow notation, natural language, "must be done before", or YAML)
- **Persistent progress**: Execution state saved to `.ralpi/progress.json` - **Parallel batching**: Independent tasks in each batch run concurrently, round-robin across configured models
- **Persistent progress**: Execution state saved to `.ralpi/progress.json`, supporting multiple PRDs simultaneously
- **Resume & auto-resume**: `/ralpi-resume` continues paused execution; a session reload mid-loop auto-resumes via `.ralpi/loop-active.json`
- **Reflection system**: Each task produces a reflection for downstream tasks - **Reflection system**: Each task produces a reflection for downstream tasks
- **Retry with backoff**: Failed tasks retry with exponential backoff - **Phased plans**: `## Phase N — Title` sections add implicit phase-boundary dependencies
- **Multiple formats**: Supports simple checkboxes, and YAML - **Model failover**: Unreachable providers cycle to the next model in the list before a task fails
- **Tool usage tracking**: Detects and reports tool usage (read, write, edit, bash) from task execution - **Auto-commit / auto-review loop**: Optional per-task commit and review-gated re-execution until pass
- **Configurable timeouts**: Task-level timeouts via meta blocks, with global fallback - **Worktree isolation**: Parallel tasks run in separate git worktrees so they can't stomp each other, with batch-level merge-conflict resolution
- **Session saving**: Saves full task output for expandable session review - **Multiple formats**: Fio README (numbered + dependencies), phased, simple checkboxes, and YAML
- **Resume auto-discovery**: Automatically finds and resumes interrupted execution - **Tool usage tracking**: Reports read/write/edit/bash usage from task execution
- **Configurable timeouts**: Task-level timeouts (inline, meta block, or YAML) with global fallback
## Usage ## Usage
``` ```
/ralpi [task-file] # Execute all tasks /ralpi [task-file] # No args → show plan for README.md; path arg → run tasks
/ralpi plan # Alias to /task-manager to plan new tasks /ralpi-run [task-file] # Execute tasks from a task file
/ralpi resume # Resume paused execution /ralpi-plan [prompt] # Open the Task Manager to plan tasks
/ralpi reset # Reset progress and .ralpi directory - does not modify PRD /ralpi-resume [task-file] # Resume paused/interrupted execution
/ralpi-reset [task-file] # Reset execution progress — does not modify the PRD
``` ```
### Highly recommended to use the task-manager prompt for prd construction, it's output pairs perfectly - /task-manager or /ralpi plan `/ralpi` with no arguments shows the execution plan for the default task file. When the first token looks like a path (`@path`, `./path`, `path/file.md`, `.yaml`, etc.) it routes to `/ralpi-run`. Everything else is handled by the dedicated dash commands above (the old `/ralpi plan|resume|reset` subcommand syntax is gone).
> The task-manager prompt (`/ralpi-plan`) pairs perfectly with ralpi's task file formats — use it for PRD construction.
## Tasks ## Tasks
### Simple Checkbox Format ### Simple Checkbox Format
```markdown ```markdown
- [ ] 01: Setup project structure - [ ] Setup project structure
- [ ] 02: Implement auth - [ ] Implement auth
- [ ] 03: Build API - [ ] Build API
```
Checkbox-only files get sequential IDs (`01`, `02`, ...). Status characters: `[ ]` pending, `[x]` done, `[~]` in progress, `[!]` failed, `[-]` skipped.
### Fio Format (numbered tasks + dependencies)
```markdown
# Build a web application
## Tasks
- [ ] 01 — Setup project structure
- [ ] 02 — Implement auth
- [ ] 03 — Build API
## Dependencies
01 -> 02, 03
``` ```
### YAML Format ### YAML Format
@@ -72,40 +96,98 @@ Use lettered sub-tasks when you discover mid-stream that a step needs to be
split. They let you preserve sibling numbering (`01`, `02`, `03`, ...) while split. They let you preserve sibling numbering (`01`, `02`, `03`, ...) while
adding granularity between two existing steps. adding granularity between two existing steps.
## Phases
`## Phase N — Title` headings group tasks into phases and add an implicit
dependency from the first task of each phase to the last task of the
previous one, so phases always run in order:
```markdown
## Phase 1 — Push-to-Talk MVP
- [ ] 01 — Voice capture
- [ ] 02 — Transmission
## Phase 2 — Group Chat
- [ ] 03 — Channels
- [ ] 04 — Presence
```
## Dependencies ## Dependencies
Dependency lines live in a `## Dependencies` section (or a plain
`Dependencies` heading). Multiple formats are supported and can be mixed.
### Arrow Notation (recommended) ### Arrow Notation (recommended)
```
1 -> 2,3,4 1 -> 2,3,4
5 -> 6 5 -> 6
This means: "Task 1 must complete before tasks 2, 3, and 4 can start." ```
"Task 1 must complete before tasks 2, 3, and 4 can start." Also supports
chains (`03 -> 04 -> 05`) and multi-prereq sources (`05, 07, 08 -> 13`).
### Natural Language ### Natural Language
```
13 depends on 17, 18, 19, 20 13 depends on 17, 18, 19, 20
14 depends on 13, 15, 16 14 depends on 13, 15, 16
22, 23, 24 depend on 21
```
This means: "Task 13 depends on tasks 17, 18, 19, and 20." "Task 13 depends on tasks 17, 18, 19, and 20." `also depends on` is accepted.
### Parallel Groups (informational only) ### "must be done before"
1, 2, 3, 4 can be done in parallel ```
21 must be done before 22, 23, 24
02, 03 must be done before 04
```
### Parallel Groups
```
1, 2, 3, 4 can be done in parallel (Play Store prep)
5, 6, 7, 8 can be done in parallel 5, 6, 7, 8 can be done in parallel
```
Note: These lines are ignored by the parser. Use explicit dependencies to control execution order. Tasks listed in a parallel group are allowed to run concurrently. Group
declarations imply no cross-group dependencies, and intra-group
dependencies are still respected — group-aware batching produces a plan
where tasks from any group run as soon as their dependencies are
satisfied.
## Configuration ## Configuration
### Task-Level Timeout ### Task-Level Timeout
You can set a timeout for individual tasks using a meta block in the task file: Timeouts can be set inline on the task line, as an inline comment, via a
meta block in the Dependencies section, or in YAML:
```markdown ```markdown
- [ ] 01: Setup project structure - [ ] 01 Setup project structure timeout: 10m
timeout: 10m - [ ] 02 — Implement auth # timeout=30s
``` ```
Supported formats: `10m` (minutes), `600s` (seconds), `3600000` (milliseconds) ```markdown
## Dependencies
01 -> 02
01 [timeout] = 10m
```
```yaml
tasks:
- id: "01"
title: Setup project structure
timeout: 15m
```
Supported units: `m` / `min` (minutes), `s` (seconds), `ms` (milliseconds).
Bare numbers default to minutes; in YAML, numeric values ≥ 1000 are treated
as milliseconds.
### Config files ### Config files
@@ -114,23 +196,31 @@ Supported formats: `10m` (minutes), `600s` (seconds), `3600000` (milliseconds)
| **Global** | `~/.pi/ralpi/config.yaml` | | **Global** | `~/.pi/ralpi/config.yaml` |
| **Project** | `./.ralpi/config.yaml` | | **Project** | `./.ralpi/config.yaml` |
Project config overrides global, which overrides defaults. Keys set
explicitly in YAML skip the corresponding loop-startup prompt.
```yaml ```yaml
execution: execution:
maxParallel: 3 # ralpi-level concurrency only maxParallel: 3 # ralpi-level concurrency only (0 = unlimited)
models: # round-robin in <provider>/<model> format models: # round-robin for parallel tasks, <provider>/<model>
- google/gemini-3.5-flash # 1st and 3rd task in parallel - anthropic/claude-sonnet-4
- openai/gpt-5.5 # 2nd task in parallel - openai/gpt-4o
autoCommit: true # commit after each task (mandated when autoReview is on; standalone toggle when off) autoCommit: true # commit after each task (mandated when autoReview is on)
autoReview: false # commit → review → loop on fail → merge on pass autoReview: false # commit → review → loop on fail → merge on pass
implModel: "" # model for task impl (sequential mode, empty = inherit parent) saveReviews: false # persist full review output to .ralpi/reviews/ (only with autoReview)
maxReviewRetries: 2 # re-executions on a 'fail' verdict before giving up
reviewBlockOnFail: false # true = mark task failed after retries exhausted instead of merging
implModel: "" # model for task impl (empty = inherit parent session model)
commitModel: "" # model for commit sessions (empty = inherit task model) commitModel: "" # model for commit sessions (empty = inherit task model)
reviewModel: "" # model for review sessions (empty = inherit task model) reviewModel: "" # model for review sessions (empty = inherit task model)
timeoutMs: 0 # per-task timeout in ms (0 = inherit Pi's defaults) timeoutMs: 0 # per-task timeout in ms (0 = inherit Pi's defaults)
commitTimeoutMs: 60000 # timeout for auto-commit agent sessions commitTimeoutMs: 0 # timeout for auto-commit agent sessions (0 = inherit)
reviewTimeoutMs: 120000 # timeout for auto-review agent sessions reviewTimeoutMs: 0 # timeout for auto-review agent sessions (0 = inherit)
loopTimeoutMs: 0 # max total loop duration in ms (0 = no limit) loopTimeoutMs: 0 # max total loop duration in ms (0 = no limit; checked between batches)
worktrees: parallel # "never" | "parallel" (default) | "always" — git worktree isolation
prompts: prompts:
projectContext: "Additional context for all tasks" projectContext: "Additional context for all tasks"
reflectionPrompt: "" # custom suffix for reflection extraction
``` ```
> `execution.models` uses slot-aware round-robin: with 3 models and 2 concurrent > `execution.models` uses slot-aware round-robin: with 3 models and 2 concurrent
@@ -140,8 +230,9 @@ prompts:
> **Automatic failover**: if a provider/API is unreachable (rate limit, 503, etc.), > **Automatic failover**: if a provider/API is unreachable (rate limit, 503, etc.),
> the task automatically cycles to the next model in the list without counting it > the task automatically cycles to the next model in the list without counting it
> as a task failure. Each model is tried once before the task is marked as failed. > as a task failure. Each model is tried once before the task is marked as failed.
> **NOTE**: this is only used in parallel execution, in sequential mode the > **NOTE**: model lists are only used in parallel execution. In sequential mode
> parent pi session's model is used > (or parallel mode with no `models` list) the parent pi session's model is used,
> unless `implModel` is set.
#### Auto-review and Auto-commit #### Auto-review and Auto-commit
@@ -158,19 +249,27 @@ state — original work plus fixes. On pass, the changes are already
committed and the worktree merges. committed and the worktree merges.
When `autoReview` is disabled, `autoCommit` runs a follow-up commit When `autoReview` is disabled, `autoCommit` runs a follow-up commit
agent after each task with no review. Both options can be overridden at agent after each task with no review. With `autoReview` on, the user is
loop startup via a selection prompt (config YAML values are honored also asked whether to persist full review output to
without prompting when set explicitly). `.ralpi/reviews/<prdKey>/<task-id>.json` (`saveReviews` — this is what
enables review feedback recovery when resuming interrupted loops). Both
options can be overridden at loop startup via a selection prompt (config
YAML values are honored without prompting when set explicitly).
`commitModel` and `reviewModel` accept `<provider>/<model>` strings (e.g. `commitModel` and `reviewModel` accept `<provider>/<model>` strings (e.g.
`anthropic/claude-sonnet-4`) resolved via the model registry. When empty, the `anthropic/claude-sonnet-4`) resolved via the model registry. When empty, the
task's model is inherited. `implModel` sets the model for task implementation task's model is inherited. `implModel` sets the model for task implementation
in sequential mode (overridden by `execution.models` round-robin in parallel (used whenever no round-robin model is assigned — sequential mode, or parallel
mode). mode with an empty `models` list; overridden by `execution.models` round-robin
in parallel mode).
## State Files ## State Files
- `.ralpi/progress.json` - Execution progress ```
- `.ralpi/reflections/` - Per-task reflections .ralpi/progress.json # Execution progress (supports multiple PRDs)
- `.ralpi/prompts/` - Generated prompts .ralpi/loop-active.json # Active-loop marker used for auto-resume after a reload
- `.ralpi/sessions/` - Full task output for review .ralpi/reflections/ # Per-task reflections
.ralpi/reviews/<prdKey>/ # Full review output (when saveReviews is on)
.ralpi/prompts/ # Generated prompts (timestamped, for debugging)
.ralpi/config.yaml # Project-level config (optional)
```

346
index.ts
View File

@@ -1,5 +1,6 @@
import * as fs from "node:fs"; import * as fs from "node:fs";
import * as path from "node:path"; import * as path from "node:path";
import { fileURLToPath } from "node:url";
import type { import type {
ExtensionAPI, ExtensionAPI,
ExtensionContext, ExtensionContext,
@@ -14,6 +15,7 @@ import {
} from "./src/dag"; } from "./src/dag";
import { ProgressTracker } from "./src/progress"; import { ProgressTracker } from "./src/progress";
import { buildPlanPrompt } from "./src/prompts"; import { buildPlanPrompt } from "./src/prompts";
import { loadTaskManagerPrompt } from "./src/task-manager-prompt";
import { formatReflections } from "./src/reflection"; import { formatReflections } from "./src/reflection";
import { verdictGlyph, verdictSummary, formatFindings } from "./src/review"; import { verdictGlyph, verdictSummary, formatFindings } from "./src/review";
import type { ReviewResult } from "./src/types"; import type { ReviewResult } from "./src/types";
@@ -21,6 +23,7 @@ import { executeBatch, type SendChatMessage } from "./src/executor";
import { import {
cleanupStaleWorktrees, cleanupStaleWorktrees,
finalizeCommittedWorktrees, finalizeCommittedWorktrees,
abortMerge,
} from "./src/worktree"; } from "./src/worktree";
import { import {
loadConfig, loadConfig,
@@ -32,6 +35,7 @@ import {
readLoopActive, readLoopActive,
findRalpiDir, findRalpiDir,
listPRDsSorted, listPRDsSorted,
countPRDResumeStats,
formatDuration, formatDuration,
} from "./src/utils"; } from "./src/utils";
@@ -197,13 +201,14 @@ async function selectLoopOptions(
/** /**
* When multiple PRD loops have progress, prompt the user to select which one * When multiple PRD loops have progress, prompt the user to select which one
* to resume. Returns the selected PRD key and sourcePath. * to act on. Returns the selected PRD key and sourcePath.
* If only one PRD exists, returns it without prompting. * If only one PRD exists, returns it without prompting.
* Returns null if no PRDs exist. * Returns null if no PRDs exist.
*/ */
async function selectPRDToResume( async function selectPRD(
ctx: ExtensionContext, ctx: ExtensionContext,
found: NonNullable<ReturnType<typeof findProgressFile>>, found: NonNullable<ReturnType<typeof findProgressFile>>,
prompt: string,
): Promise<{ prdKey: string; sourcePath: string } | null> { ): Promise<{ prdKey: string; sourcePath: string } | null> {
const prds = listPRDsSorted(found.state); const prds = listPRDsSorted(found.state);
if (prds.length === 0) return null; if (prds.length === 0) return null;
@@ -213,25 +218,20 @@ async function selectPRDToResume(
// Multiple PRDs — show selection sorted by most recent first // Multiple PRDs — show selection sorted by most recent first
const options = prds.map((entry) => { const options = prds.map((entry) => {
const tasks = entry.prd.tasks; // Total/completed must come from the parsed PRD file, not just the
const total = Object.keys(tasks).length; // progress map: the tracker only records TOUCHED tasks (started/
const completed = Object.values(tasks).filter( // completed/failed), so never-started tasks would be silently missing
(t) => t.status === "completed", // from a naive Object.keys() count and the totals would under-report.
).length; const { total, completed, failed } = countPRDResumeStats(
const failed = Object.values(tasks).filter( entry.prd,
(t) => t.status === "failed", entry.prd.sourcePath,
).length; );
const relPath = path.relative(ctx.cwd, entry.prd.sourcePath); const relPath = path.relative(ctx.cwd, entry.prd.sourcePath);
const updated = new Date(entry.prd.lastUpdatedAt).toLocaleString(); const updated = new Date(entry.prd.lastUpdatedAt).toLocaleString();
return `${relPath}${completed}/${total} done${ return `${relPath}${completed}/${total} done${failed ? `, ${failed} failed` : ""} · ${updated}`;
failed ? `, ${failed} failed` : ""
} · ${updated}`;
}); });
const selected = await ctx.ui.select( const selected = await ctx.ui.select(prompt, options);
"Multiple loops found. Which to resume?",
options,
);
if (!selected) return null; if (!selected) return null;
const idx = options.indexOf(selected); const idx = options.indexOf(selected);
@@ -255,6 +255,23 @@ async function executePlanBatches(
projectDir?: string, projectDir?: string,
isResume?: boolean, isResume?: boolean,
): Promise<void> { ): Promise<void> {
// Refresh the model registry so the host reloads models.json before we
// resolve the round-robin model pool. The registry snapshot is captured at
// host startup and only reloaded here; a long-running host would otherwise
// skip providers added to models.json after it booted (e.g. "strix").
// Best-effort: a failed refresh shouldn't block execution — the pool just
// resolves against the existing snapshot.
try {
await ctx.modelRegistry?.refresh();
} catch (error) {
ctx.ui.notify(
`ralpi: model registry refresh failed — continuing with existing snapshot: ${
error instanceof Error ? error.message : String(error)
}`,
"warning",
);
}
// Write loop-active marker so a session reload can detect an interrupted // Write loop-active marker so a session reload can detect an interrupted
// loop and resume it (in-process agent sessions die on reload — the marker // loop and resume it (in-process agent sessions die on reload — the marker
// + progress.json in_progress tasks are the signal to re-run them). // + progress.json in_progress tasks are the signal to re-run them).
@@ -566,27 +583,9 @@ export default function ralpiLoopExtension(pi: ExtensionAPI): void {
t.status === "in_progress" ? [id] : [], t.status === "in_progress" ? [id] : [],
); );
if (inProgressIds.length === 0) {
// Nothing was mid-flight — loop either finished cleanly between
// the reload landing and this handler running, or was stopped
// between tasks. Clean up the stale marker and bail.
ctx.ui.notify(
"ralpi loop has no in-progress task to resume — marking complete.",
"info",
);
deleteLoopActive(projectDir);
return;
}
const taskCount = loopState.taskIds.length;
ctx.ui.notify(
`ralpi loop was interrupted by reload with ${inProgressIds.length} in-progress task(s). ` +
`Resuming execution (${taskCount} tasks, ${loopState.mode} mode)...`,
"info",
);
// Build the sendProgress wrapper so resumed task messages render the // Build the sendProgress wrapper so resumed task messages render the
// same expandable tool-call tree as an interactive run. // same expandable tool-call tree as an interactive run. Defined before
// the finalize path below so it can report self-healed merges.
const sendProgress: SendChatMessage = ( const sendProgress: SendChatMessage = (
content: string, content: string,
meta?: { meta?: {
@@ -610,6 +609,131 @@ export default function ralpiLoopExtension(pi: ExtensionAPI): void {
}); });
}; };
if (inProgressIds.length === 0) {
// Nothing was mid-flight — the loop either finished cleanly between
// the reload landing and this handler running, or was stopped
// between tasks. Either way, committed worktree branches from an
// interrupted loop may still be unmerged (e.g. a prior resume
// attempt reset tasks to pending before it was itself interrupted).
// Finalize those first so committed code lands in the workspace,
// persist the state to progress.json, update the PRD file, THEN
// clean up the stale marker.
try {
const config = loadConfig(projectDir);
// Clear any half-done merge left by an interrupted
// conflict-resolution session (it would block every merge below).
abortMerge(projectDir);
const allIds = Object.entries(initialTasks).flatMap(([id, t]) =>
t.status !== "failed" && t.status !== "pending" ? [id] : [],
);
const fin = finalizeCommittedWorktrees(
projectDir,
config.paths.stateDir,
loopState.prdKey,
allIds,
);
// Persist finalized tasks to progress.json + PRD file so the
// state is correct for subsequent /ralpi resume calls.
const stateDir = config.paths.stateDir;
const progressPath = path.join(projectDir, stateDir, "progress.json");
// Batch-update progress.json and PRD file for all finalized tasks
if (fin.finalized.length > 0) {
const progressRaw = fs.existsSync(progressPath)
? JSON.parse(fs.readFileSync(progressPath, "utf-8"))
: null;
for (const id of fin.finalized) {
sendProgress?.(
`${id} — finalized on resume (committed branch merged into main)`,
);
if (progressRaw) {
const tasks =
progressRaw.prds?.[loopState.prdKey]?.tasks ??
progressRaw.tasks;
if (tasks && tasks[id]) {
tasks[id].status = "completed";
tasks[id].completedAt = new Date().toISOString();
}
}
try {
const prdPath = loopState.taskFile;
if (fs.existsSync(prdPath)) {
updateTaskInFile(prdPath, id, "completed");
}
} catch {
// Best-effort
}
}
if (progressRaw) {
fs.writeFileSync(
progressPath,
JSON.stringify(progressRaw, null, 2),
"utf-8",
);
}
}
// ── Handle conflicted tasks ──
// Same logic as resumeLoop: reset to pending so the DAG can
// re-schedule them, keep the worktree for in-place re-run.
const conflictIds = Object.keys(fin.conflicts);
if (conflictIds.length > 0) {
const detail = conflictIds
.map((id) => `${id}: ${fin.conflicts[id].slice(0, 3).join(", ")}`)
.join("; ");
// Batch-reset all conflicted tasks to pending, then write once
const progressRaw = fs.existsSync(progressPath)
? JSON.parse(fs.readFileSync(progressPath, "utf-8"))
: null;
for (const id of conflictIds) {
if (progressRaw) {
const tasks =
progressRaw.prds?.[loopState.prdKey]?.tasks ??
progressRaw.tasks;
if (tasks && tasks[id]) {
tasks[id].status = "pending";
delete tasks[id].startedAt;
delete tasks[id].error;
}
}
try {
const prdPath = loopState.taskFile;
if (fs.existsSync(prdPath)) {
updateTaskInFile(prdPath, id, "pending");
}
} catch {
// Best-effort
}
}
if (progressRaw) {
fs.writeFileSync(
progressPath,
JSON.stringify(progressRaw, null, 2),
"utf-8",
);
}
ctx.ui.notify(
`Reset ${conflictIds.length} conflicted task(s) to pending for re-execution (${detail})`,
"info",
);
}
} catch {
// Best-effort — the marker is removed either way; the worktrees
// stay on disk for a manual /ralpi-resume.
}
ctx.ui.notify(
"ralpi loop has no in-progress task to resume — marking complete.",
"info",
);
deleteLoopActive(projectDir);
return;
}
const taskCount = loopState.taskIds.length;
ctx.ui.notify(
`ralpi loop was interrupted by reload with ${inProgressIds.length} in-progress task(s). ` +
`Resuming execution (${taskCount} tasks, ${loopState.mode} mode)...`,
"info",
);
// Load config from the project directory so model + thinking level // Load config from the project directory so model + thinking level
// resolve the same way the interactive command handler does. // resolve the same way the interactive command handler does.
const config = loadConfig(projectDir); const config = loadConfig(projectDir);
@@ -692,12 +816,18 @@ export default function ralpiLoopExtension(pi: ExtensionAPI): void {
}, },
}); });
const extensionDir = path.dirname(fileURLToPath(import.meta.url));
pi.registerCommand("ralpi-plan", { pi.registerCommand("ralpi-plan", {
description: "Open the Task Manager to plan a ralpi run", description: "Open the Task Manager to plan a ralpi run",
handler: async (args: string, ctx: ExtensionContext) => { handler: async (args: string, ctx: ExtensionContext) => {
const prompt = (args || "").trim(); // pi.sendUserMessage() sends with expandPromptTemplates: false, so it
const message = prompt ? `@task-manager\n\n${prompt}` : "@task-manager"; // would NOT expand `/task-manager` — and `@task-manager` is an
pi.sendUserMessage(message); // @-mention, not a template invocation. Load the bundled template,
// strip frontmatter, substitute $@ args ourselves, and send the
// expanded body directly.
const body = loadTaskManagerPrompt(extensionDir, args ?? "");
pi.sendUserMessage(body);
ctx.ui.notify("Opening Task Manager...", "info"); ctx.ui.notify("Opening Task Manager...", "info");
}, },
}); });
@@ -897,22 +1027,37 @@ async function resumeLoop(
// A review-gated task whose agent committed + reviewed successfully still // A review-gated task whose agent committed + reviewed successfully still
// needs a final merge into main + worktree removal to be "done". If the // needs a final merge into main + worktree removal to be "done". If the
// loop was interrupted between that commit and the merge, the task is left // loop was interrupted between that commit and the merge, the task is left
// `in_progress` with a clean, committed worktree branch. Resuming without // with a committed worktree branch. Resuming without finalizing would
// finalizing would wastefully re-run finished work. // wastefully re-run finished work — or worse, strand the committed code in
// `.ralpi/worktrees/` forever.
// //
// finalizeCommittedWorktrees merges those branches into main now; the // finalizeCommittedWorktrees runs over EVERY non-failed task, not just
// rest (dirty trees, nothing committed, conflicts) are left in_progress // `in_progress` ones: a prior interrupted resume can reset tasks to
// and reset to pending below for a real re-run. // `pending` while their worktree branch still holds committed work that
// was never merged. Only scanning in_progress tasks would silently leave
// that code out of the workspace on every resume.
//
// `pending` tasks (never started) are excluded — they never had worktrees
// created, so finalize always puts them in `rerun`, which is wasted work.
// Failed tasks keep their worktrees for inspection/re-run and are also
// deliberately excluded.
const prdKeyForFinalize = progress.getKey(); const prdKeyForFinalize = progress.getKey();
const inProgressIds = Object.entries(progress.getState().tasks) // Clear any half-done merge left in the main repo by an interrupted
.filter(([, t]) => t.status === "in_progress") // conflict-resolution session — it would block every merge below
.map(([id]) => id); // (`git merge` refuses while a merge is already in progress). No-op when
if (inProgressIds.length > 0) { // the repo isn't mid-merge.
abortMerge(projectDir);
const finalizeCandidateIds = Object.entries(
progress.getState().tasks,
).flatMap(([id, t]) =>
t.status !== "failed" && t.status !== "pending" ? [id] : [],
);
if (finalizeCandidateIds.length > 0) {
const fin = finalizeCommittedWorktrees( const fin = finalizeCommittedWorktrees(
projectDir, projectDir,
config.paths.stateDir, config.paths.stateDir,
prdKeyForFinalize, prdKeyForFinalize,
inProgressIds, finalizeCandidateIds,
); );
for (const id of fin.finalized) { for (const id of fin.finalized) {
progress.markCompleted(id, 0); progress.markCompleted(id, 0);
@@ -925,15 +1070,47 @@ async function resumeLoop(
`${id} — finalized on resume (committed branch merged into main)`, `${id} — finalized on resume (committed branch merged into main)`,
); );
} }
// ── Handle conflicted tasks ──
//
// Tasks whose committed branch could not be auto-merged (git conflicts)
// must be reset to `pending` so the DAG re-schedules them. The worktree
// is preserved — the agent re-runs in-place in the existing worktree via
// createWorktree's reuse logic. If the agent's re-run changes make the
// merge succeed on the next attempt, the loop continues normally. If the
// merge fails again, `executeBatch`'s batch-level conflict resolution
// (`resolveConflictsSession`) handles the conflict markers properly.
const conflictIds = Object.keys(fin.conflicts); const conflictIds = Object.keys(fin.conflicts);
if (conflictIds.length > 0) { if (conflictIds.length > 0) {
const detail = conflictIds const detail = conflictIds
.map((id) => `${id}: ${fin.conflicts[id].slice(0, 3).join(", ")}`) .map((id) => `${id}: ${fin.conflicts[id].slice(0, 3).join(", ")}`)
.join("; "); .join("; ");
// Batch-reset all conflicted tasks to pending, then save once.
// Directly mutate the progress state (there's no markPending method
// on ProgressTracker — markFailed would leave it as 'failed' which the
// DAG excludes). The worktree is preserved so createWorktree reuses
// it and the agent re-runs in-place.
const tasks = progress.getState().tasks;
for (const id of conflictIds) {
if (tasks[id]) {
tasks[id].status = "pending";
delete tasks[id].startedAt;
delete tasks[id].error;
}
try {
updateTaskInFile(taskFile, id, "pending");
} catch {
// Best-effort
}
}
progress.save();
sendChatMessage?.( sendChatMessage?.(
`${conflictIds.join( `${conflictIds.join(
", ", ", ",
)} — merge conflict on resume-finalize; re-running (${detail})`, )} — merge conflict on resume-finalize; reset to pending for re-execution (${detail})`,
);
ctx.ui.notify(
`Reset ${conflictIds.length} conflicted task(s) to pending for re-execution`,
"info",
); );
} }
} }
@@ -1052,7 +1229,11 @@ async function handleResume(
// When no specific task file is given, let the user select which loop // When no specific task file is given, let the user select which loop
// to resume from multiple PRDs (sorted by most recent first). // to resume from multiple PRDs (sorted by most recent first).
const selected = await selectPRDToResume(ctx, found); const selected = await selectPRD(
ctx,
found,
"Multiple loops found. Which to resume?",
);
if (!selected) { if (!selected) {
ctx.ui.notify("Resume cancelled.", "info"); ctx.ui.notify("Resume cancelled.", "info");
return; return;
@@ -1106,12 +1287,17 @@ async function handleReset(
ctx: ExtensionContext, ctx: ExtensionContext,
args: string[], args: string[],
): Promise<void> { ): Promise<void> {
let sourcePath: string;
let prdKey: string | undefined;
let progress: ProgressTracker;
if (args[0]) { if (args[0]) {
const taskFile = resolveTaskArg(args[0], ctx.cwd); const taskFile = resolveTaskArg(args[0], ctx.cwd);
const found = findProgressFile(ctx.cwd, taskFile); const found = findProgressFile(ctx.cwd, taskFile);
const projectDir = found ? path.dirname(path.dirname(found.path)) : ctx.cwd; const projectDir = found ? path.dirname(path.dirname(found.path)) : ctx.cwd;
const progress = new ProgressTracker(projectDir, taskFile, found?.prdKey); sourcePath = taskFile;
progress.reset(); prdKey = found?.prdKey;
progress = new ProgressTracker(projectDir, taskFile, prdKey);
} else { } else {
const found = findProgressFile(ctx.cwd); const found = findProgressFile(ctx.cwd);
if (!found) { if (!found) {
@@ -1122,13 +1308,53 @@ async function handleReset(
return; return;
} }
const projectDir = path.dirname(path.dirname(found.path)); const projectDir = path.dirname(path.dirname(found.path));
// Use the most recently updated PRD (first in sorted order)
const prds = listPRDsSorted(found.state); // Multiple loops may have progress — let the user select which one to
const sourcePath = // reset (sorted by most recent first), same as resume.
prds.length > 0 ? prds[0].prd.sourcePath : found.state.sourcePath; const selected = await selectPRD(
const progress = new ProgressTracker(projectDir, sourcePath); ctx,
progress.reset(); found,
"Multiple loops found. Which to reset?",
);
if (!selected) {
ctx.ui.notify("Reset cancelled.", "info");
return;
}
sourcePath = selected.sourcePath;
prdKey = selected.prdKey;
progress = new ProgressTracker(projectDir, sourcePath, prdKey);
} }
// Ask whether to also clear the progress markers (checkboxes/status) in the
// source PRD/README file itself, not just .ralpi/progress.json.
const taskIds = Object.keys(progress.getState().tasks);
if (taskIds.length > 0) {
const choice = await ctx.ui.select(
`Also reset progress markers in the source file (${path.basename(
sourcePath,
)})?`,
[
"Yes — clear checkboxes/status markers in the source PRD/README",
"No — only reset .ralpi/progress.json",
],
);
if (choice === undefined) {
ctx.ui.notify("Reset cancelled.", "info");
return;
}
if (choice.startsWith("Yes")) {
progress.reset();
for (const id of taskIds) {
updateTaskInFile(sourcePath, id, "pending");
}
ctx.ui.notify(
`Progress reset — cleared ${taskIds.length} task marker(s) in ${path.basename(sourcePath)}.`,
"info",
);
return;
}
}
progress.reset();
ctx.ui.notify("Progress reset. All task statuses cleared.", "info"); ctx.ui.notify("Progress reset. All task statuses cleared.", "info");
} }

View File

@@ -1,6 +1,6 @@
{ {
"name": "@mikefreno/ralpi", "name": "@mikefreno/ralpi",
"version": "0.3.0", "version": "0.5.0",
"description": "Execute tasks from task files/PRD's using DAG-based dependency resolution with persistent progress tracking", "description": "Execute tasks from task files/PRD's using DAG-based dependency resolution with persistent progress tracking",
"keywords": [ "keywords": [
"pi-package", "pi-package",

View File

@@ -1,38 +1,44 @@
--- ---
description: Executes individual tasks from ralpi task files using DAG-based dependency resolution, with progress tracking and reflection support description: Execute tasks from ralpi task files / PRDs using DAG-based dependency resolution, with persistent progress tracking and reflection support
--- ---
# ralpi-task # ralpi-task
Execute a single task from a ralpi task file. Execute tasks from a ralpi task file (checkbox, Fio, phased, or YAML) using DAG-based dependency resolution, with persistent progress tracking and reflection support.
## When to Use ## When to Use
- User asks to execute a specific task from a task file - User asks to execute a task file, PRD, or task list (e.g. "run the tasks", "execute the plan")
- User provides a task ID and wants to run it - User wants to run a full ralpi loop from a task file in the project
- User wants to run the next task in sequence - User wants to resume an interrupted or paused ralpi run
- User wants to plan new tasks with the task-manager prompt
## Usage ## Usage
``` ```
/ralpi run [task-file] # Run all tasks /ralpi [task-file] # No args → show plan; path arg → run all tasks
/ralpi next [task-file] # Run next batch /ralpi-run [task-file] # Run all tasks (auto-resumes if progress exists)
/ralpi status [task-file] # Check progress /ralpi-resume [task-file] # Resume paused/interrupted execution
/ralpi-plan [prompt] # Open the Task Manager to plan tasks
``` ```
Note: ralpi runs whole task plans — there is no single-task or next-batch subcommand. To execute only part of a plan, edit the task file and remove/adjust the tasks first.
## Task File Location ## Task File Location
Default: `README.md` in current directory. Can be overridden with explicit path. Default: `README.md` in the current directory. Can be overridden with an explicit path (`@path`, `./path`, `*.md`, `*.yaml`, `*.yml`).
## Reflection Format ## Reflection Format
After completing a task, include: After completing a task, the task agent ends its response with a reflection block, which the extension parses and passes to downstream tasks:
``` ```
## REFLECTION ## REFLECTION
SUMMARY: [what was done] SUMMARY: [1-2 sentence description of what was accomplished]
FILES: [files changed] FILES: [comma-separated list of files created or modified]
LEARNINGS: LEARNINGS:
- [key learning] - [key decision, pattern, or architectural choice]
BLOCKERS: [issues or 'none'] - [important API or interface details]
- [anything downstream tasks need to know]
BLOCKERS: [any unresolved issues, or 'none']
``` ```

View File

@@ -35,6 +35,7 @@ import {
abortMerge, abortMerge,
hasMergeConflicts, hasMergeConflicts,
completeMerge, completeMerge,
worktreeHasPreservableWork,
type WorktreeHandle, type WorktreeHandle,
type MergeResult, type MergeResult,
} from "./worktree"; } from "./worktree";
@@ -1248,7 +1249,7 @@ async function executeTask(
}`, }`,
"error", "error",
); );
if (wt) removeWorktree(projectDir, wt); cleanupFailedWorktree(projectDir, wt, task, sendChatMessage);
roundRobin?.release(task.id); roundRobin?.release(task.id);
return; return;
} catch (error) { } catch (error) {
@@ -1264,7 +1265,7 @@ async function executeTask(
} }
sendChatMessage?.(`${task.id} · ${task.title}${errorMsg}`); sendChatMessage?.(`${task.id} · ${task.title}${errorMsg}`);
ctx.ui.notify(`Task ${task.id} failed: ${errorMsg}`, "error"); ctx.ui.notify(`Task ${task.id} failed: ${errorMsg}`, "error");
if (wt) removeWorktree(projectDir, wt); cleanupFailedWorktree(projectDir, wt, task, sendChatMessage);
return; return;
} }
} }
@@ -1280,11 +1281,37 @@ async function executeTask(
`Task ${task.id} failed: all configured models exhausted`, `Task ${task.id} failed: all configured models exhausted`,
"error", "error",
); );
if (wt) removeWorktree(projectDir, wt); cleanupFailedWorktree(projectDir, wt, task, sendChatMessage);
} }
// ─── Save Reflection to File ──────────────────────────────────────────────── // ─── Save Reflection to File ────────────────────────────────────────────────
/**
* Remove a task worktree after a failure UNLESS it still holds recoverable
* work (commits ahead of main, or uncommitted changes).
*
* `removeWorktree` force-deletes the worktree's branch, which makes any
* commits the agent made before failing/timing out unreachable — real code
* loss. A preserved worktree is instead picked up on the next resume:
* resume-finalize merges committed work into main, or the task re-runs in
* place and the agent continues from where it stopped.
*/
function cleanupFailedWorktree(
projectDir: string,
wt: WorktreeHandle | null,
task: Task,
sendChatMessage?: SendChatMessage,
): void {
if (!wt) return;
if (worktreeHasPreservableWork(projectDir, wt)) {
sendChatMessage?.(
`~ ${task.id} · ${task.title} — task failed but worktree preserved (${wt.branch}); committed work will be merged on resume`,
);
return;
}
removeWorktree(projectDir, wt);
}
function saveReflectionToFile( function saveReflectionToFile(
sourceDir: string, sourceDir: string,
config: RalpiConfig, config: RalpiConfig,
@@ -1407,7 +1434,7 @@ async function runFollowUpSession(
undefined, undefined,
model, model,
config.thinkingLevel, config.thinkingLevel,
true, // noSkills — follow-up sessions don't need the skills catalog false, // noSkills=false — follow-up sessions load skills too
(ctx.modelRegistry as any).runtime as ModelRuntime, (ctx.modelRegistry as any).runtime as ModelRuntime,
); );

View File

@@ -142,6 +142,33 @@ export class ProgressTracker {
/** Save current state to disk */ /** Save current state to disk */
save(): void { save(): void {
// Merge into the freshest on-disk state instead of writing the
// construction-time snapshot verbatim. Each ProgressTracker instance
// (one per PRD loop) snapshots the WHOLE state at construction; when
// two loops run concurrently in one project, saving a stale snapshot
// would silently revert the OTHER loop's task status changes — tasks
// get wrongly written back to "pending" while their worktrees carry
// real work, stranding it on the next resume.
let disk: ProgressState | null = null;
try {
if (fs.existsSync(this.statePath)) {
const raw = fs.readFileSync(this.statePath, "utf-8");
disk = JSON.parse(raw) as ProgressState;
}
} catch {
disk = null;
}
if (disk && disk.prds) {
// Keep THIS tracker's in-memory PRD (its own tasks are the source
// of truth — all status mutations happened on it), but adopt the
// on-disk entries for OTHER PRDs instead of writing the stale
// construction-time snapshot over them.
const mine = this.getPRD();
this.state = disk;
this.state.prds ??= {};
this.state.prds[this.prdKey] = mine;
}
const prd = this.getPRD(); const prd = this.getPRD();
prd.lastUpdatedAt = new Date().toISOString(); prd.lastUpdatedAt = new Date().toISOString();
// Sync legacy flat fields with current PRD for backward compat // Sync legacy flat fields with current PRD for backward compat

104
src/task-manager-prompt.ts Normal file
View File

@@ -0,0 +1,104 @@
import * as fs from "node:fs";
import * as path from "node:path";
import { stripFrontmatter } from "@earendil-works/pi-coding-agent";
const TEMPLATE_REL = path.join("prompts", "task-manager.md");
/**
* Parse command arguments respecting quoted strings (bash-style).
* Ported from pi's core/prompt-templates.js so the task-manager template
* receives the same arg-splitting a real `/task-manager` invocation would.
*/
function parseCommandArgs(argsString: string): string[] {
const args: string[] = [];
let current = "";
let inQuote: string | null = null;
for (let i = 0; i < argsString.length; i++) {
const char = argsString[i];
if (inQuote) {
if (char === inQuote) {
inQuote = null;
} else {
current += char;
}
} else if (char === '"' || char === "'") {
inQuote = char;
} else if (/\s/.test(char)) {
if (current) {
args.push(current);
current = "";
}
} else {
current += char;
}
}
if (current) args.push(current);
return args;
}
/**
* Substitute argument placeholders in template content.
* Faithful port of pi's substituteArgs (core/prompt-templates.js):
* - $1, $2, ... positional args
* - $@ / $ARGUMENTS all args joined
* - ${N:-default} positional N with default when missing/empty
* - ${@:-default} all args with default when empty
* - ${@:N} / ${@:N:L} bash-style slicing
*
* Replacement runs once over the template only; argument/default values
* containing patterns like $1 or $@ are NOT recursively substituted.
*/
function substituteArgs(content: string, args: string[]): string {
const allArgs = args.join(" ");
return content.replace(
/\$\{(\d+|ARGUMENTS|@):-([^}]*)\}|\$\{@:(\d+)(?::(\d+))?\}|\$(ARGUMENTS|@|\d+)/g,
(_match, defaultTarget, defaultValue, sliceStart, sliceLength, simple) => {
if (defaultTarget) {
const value =
defaultTarget === "@" || defaultTarget === "ARGUMENTS"
? allArgs
: args[parseInt(defaultTarget, 10) - 1];
return value ? value : defaultValue;
}
if (sliceStart) {
let start = parseInt(sliceStart, 10) - 1; // 1-indexed → 0-indexed
if (start < 0) start = 0;
if (sliceLength) {
const length = parseInt(sliceLength, 10);
return args.slice(start, start + length).join(" ");
}
return args.slice(start).join(" ");
}
if (simple === "ARGUMENTS" || simple === "@") {
return allArgs;
}
const index = parseInt(simple, 10) - 1;
return args[index] ?? "";
},
);
}
/**
* Load and expand the task-manager prompt template bundled with the extension.
*
* `pi.sendUserMessage()` sends with `expandPromptTemplates: false`, so it will
* NOT expand a `/task-manager` invocation — and `@task-manager` is an
* @-mention, not a template invocation anyway. We therefore read the
* template ourselves, strip its frontmatter, substitute args ($@ etc.), and
* return the fully-expanded prompt body ready to send as a user message.
*
* @param extensionDir Absolute path to the extension root (where index.ts
* lives), used to locate `prompts/task-manager.md`.
* @param argsString Raw argument string from the slash command (may be "").
* @throws if the template file is missing or unreadable.
*/
export function loadTaskManagerPrompt(
extensionDir: string,
argsString: string,
): string {
const templatePath = path.join(extensionDir, TEMPLATE_REL);
const raw = fs.readFileSync(templatePath, "utf-8");
const body = stripFrontmatter(raw);
const args = parseCommandArgs(argsString);
return substituteArgs(body, args).trim();
}

View File

@@ -7,6 +7,7 @@ import type {
ToolUsage, ToolUsage,
} from "./types"; } from "./types";
import { DEFAULT_CONFIG } from "./types"; import { DEFAULT_CONFIG } from "./types";
import { parseTaskFile } from "./parser";
import type { AgentSessionEvent } from "@earendil-works/pi-coding-agent"; import type { AgentSessionEvent } from "@earendil-works/pi-coding-agent";
import { import {
createAgentSession, createAgentSession,
@@ -198,6 +199,49 @@ export function listPRDsSorted(
return entries; return entries;
} }
export interface PRDResumeSummary {
total: number;
completed: number;
failed: number;
}
/**
* Count tasks for the resume-selection display.
*
* The progress tracker only records tasks that were TOUCHED (started,
* completed, or failed) — never-started tasks are absent from `prd.tasks`,
* so a naive Object.keys(prd.tasks).length under-reports the real total.
* The true total comes from parsing the PRD source file. Completed counts
* both progress-marked completions and PRD checkbox completions (a task
* checked off in the file is done even if the loop was interrupted before
* markCompleted), deduped by task id. Falls back to touched-task counts
* when the source file is missing or unparseable.
*/
export function countPRDResumeStats(
prd: PRDProgress,
sourcePath: string,
): PRDResumeSummary {
const touched = Object.entries(prd.tasks);
const failed = touched.filter(([, t]) => t.status === "failed").length;
const completedIds = new Set(
touched.filter(([, t]) => t.status === "completed").map(([id]) => id),
);
let total: number;
try {
const project = parseTaskFile(sourcePath);
total = project.tasks.length;
for (const task of project.tasks) {
if (task.status === "completed") completedIds.add(task.id);
}
} catch {
// PRD file missing/unparseable — fall back to touched-task counts
total = touched.length;
}
return { total, completed: completedIds.size, failed };
}
// ─── Model Resolution ─────────────────────────────────────────────────────── // ─── Model Resolution ───────────────────────────────────────────────────────
/** /**
@@ -552,14 +596,17 @@ export async function runAgentSession(
} = {}; } = {};
try { try {
// Loop sessions load the full normal pi context: extensions (so all
// extension-provided tools register), skills, and project context
// (AGENTS.md / CLAUDE.md)
const loader = new DefaultResourceLoader({ const loader = new DefaultResourceLoader({
cwd, cwd,
agentDir: getAgentDir(), agentDir: getAgentDir(),
noExtensions: true,
noSkills, noSkills,
noPromptTemplates: true, noPromptTemplates: true,
noThemes: true, noThemes: true,
noContextFiles: true, noExtensions: false,
noContextFiles: false,
}); });
await loader.reload(); await loader.reload();
@@ -569,7 +616,7 @@ export async function runAgentSession(
resourceLoader: loader, resourceLoader: loader,
settingsManager: SettingsManager.create(cwd, getAgentDir()), settingsManager: SettingsManager.create(cwd, getAgentDir()),
modelRuntime, modelRuntime,
tools: ["read", "bash", "edit", "write", "grep", "find", "ls"], // No `tools` allowlist: matches a normal pi session's tool set.
model: model as any, model: model as any,
thinkingLevel: thinkingLevel as any, thinkingLevel: thinkingLevel as any,
}); });
@@ -673,6 +720,8 @@ function extractAssistantText(content: unknown): string {
/** /**
* Check if there are any uncommitted changes in the git repository. * Check if there are any uncommitted changes in the git repository.
* Includes untracked files — a new file created by a task agent is work
* that still needs committing.
*/ */
export function hasUncommittedChanges(projectDir: string): boolean { export function hasUncommittedChanges(projectDir: string): boolean {
const { execSync } = require("node:child_process"); const { execSync } = require("node:child_process");
@@ -687,6 +736,32 @@ export function hasUncommittedChanges(projectDir: string): boolean {
} }
} }
/**
* Check for uncommitted changes to TRACKED files only, ignoring untracked
* (`??`) entries.
*
* Untracked files never block a merge, so a worktree whose task work is
* fully committed is "done" even when it carries stray untracked files
* (scratch files, build artifacts, files created but deliberately left out
* of the commit). Resume-finalize uses this to decide whether a task's
* committed branch should be merged into main: counting `??` entries there
* would strand committed code in `.ralpi/worktrees/` forever.
*/
export function hasTrackedUncommittedChanges(projectDir: string): boolean {
const { execSync } = require("node:child_process");
try {
const output = execSync("git status --porcelain", {
cwd: projectDir,
encoding: "utf-8",
}).trim();
return output
.split("\n")
.some((line: string) => line.length > 0 && !line.startsWith("??"));
} catch {
return false;
}
}
/** /**
* Get the current git status in porcelain format. * Get the current git status in porcelain format.
* Includes untracked files, which `git diff` alone would miss. * Includes untracked files, which `git diff` alone would miss.

View File

@@ -1,6 +1,10 @@
import * as fs from "node:fs"; import * as fs from "node:fs";
import * as path from "node:path"; import * as path from "node:path";
import { ensureDir, hasUncommittedChanges } from "./utils"; import {
ensureDir,
hasUncommittedChanges,
hasTrackedUncommittedChanges,
} from "./utils";
// ─── Types ─────────────────────────────────────────────────────────────────── // ─── Types ───────────────────────────────────────────────────────────────────
@@ -83,6 +87,28 @@ export function getCurrentBranch(dir: string): string | null {
return git("rev-parse --abbrev-ref HEAD", dir); return git("rev-parse --abbrev-ref HEAD", dir);
} }
/**
* Canonicalize a directory path, resolving symlinks.
*
* `git worktree list --porcelain` emits REAL paths (symlinks resolved,
* e.g. `/private/tmp/...` for `/tmp/...` on macOS), while `path.join` on a
* caller-supplied path keeps the literal spelling. Comparing the two
* verbatim silently fails — resume then can't see an existing worktree,
* `createWorktree` falls through to a fresh `worktree add` that fails
* because the directory already exists, returns null, and the task agent
* ends up running in the MAIN repo with no worktree merge at all.
*
* All worktree path computation and porcelain comparisons go through this
* so literal vs real paths can never diverge.
*/
function canonicalDir(dir: string): string {
try {
return fs.realpathSync(dir);
} catch {
return path.resolve(dir);
}
}
// ─── Worktree Lifecycle ────────────────────────────────────────────────────── // ─── Worktree Lifecycle ──────────────────────────────────────────────────────
/** /**
@@ -152,6 +178,9 @@ export function createWorktree(
baseRef?: string, baseRef?: string,
taskTitle?: string, taskTitle?: string,
): WorktreeHandle | null { ): WorktreeHandle | null {
// Canonicalize FIRST: every path below (worktree dir, porcelain
// comparisons, branch refs) must share one spelling of the repo path.
mainDir = canonicalDir(mainDir);
if (!isGitRepo(mainDir)) return null; if (!isGitRepo(mainDir)) return null;
const safeId = safeBranchSuffix(taskId); const safeId = safeBranchSuffix(taskId);
@@ -172,6 +201,23 @@ export function createWorktree(
if (existing && existing.includes(`worktree ${wtDir}`)) { if (existing && existing.includes(`worktree ${wtDir}`)) {
// The worktree is registered — sanity-check it's a valid checkout. // The worktree is registered — sanity-check it's a valid checkout.
if (getGitHead(wtDir)) { if (getGitHead(wtDir)) {
// Return the branch the worktree is ACTUALLY checked out on, NOT the
// slug recomputed from the (possibly changed) task title. Mismatch
// happens routinely on resume: the task agent may have created its
// own feature branch (e.g. `proctored-exam-delivery-mode-10-exam-...`)
// once it saw the convention in the git log, the title may have been
// edited between runs, or an older ralpi version used a different
// naming scheme. Returning the slug here makes `git merge <slug>`
// fail with "not something we can merge" because no such ref exists —
// exactly the spurious merge-conflict we see on resumes.
const actual = getCurrentBranch(wtDir);
if (actual && actual !== "HEAD" && actual !== "detached") {
return { dir: wtDir, branch: actual, mainDir };
}
// Detached-HEAD worktree (e.g. left by a prior `--detach` fallback).
// The slug ref doesn't exist as a branch — create one matching the
// slug from the worktree's current HEAD so the merge step resolves.
git(`branch "${branch}" HEAD`, mainDir);
return { dir: wtDir, branch, mainDir }; return { dir: wtDir, branch, mainDir };
} }
// Registered but broken (dir gone / checkout corrupt) — drop its // Registered but broken (dir gone / checkout corrupt) — drop its
@@ -340,7 +386,7 @@ export function cleanupStaleWorktrees(
// When a prdKey is given, narrow to that PRD's subdir so concurrent // When a prdKey is given, narrow to that PRD's subdir so concurrent
// loops (other PRDs) are not disturbed. // loops (other PRDs) are not disturbed.
const managedRoot = path.resolve( const managedRoot = path.resolve(
mainDir, canonicalDir(mainDir),
stateDir, stateDir,
"worktrees", "worktrees",
...(prdKey ? [prdKey] : []), ...(prdKey ? [prdKey] : []),
@@ -391,20 +437,27 @@ export interface FinalizeResult {
} }
/** /**
* Finalize in-progress tasks whose worktrees already hold committed, clean * Finalize worktrees that already hold committed, clean work that was never
* work that was never merged into main (typically because the loop was * merged into main (typically because the loop was interrupted between the
* interrupted between the task commit and the merge/finalize step). * task commit and the merge/finalize step).
* *
* For each task ID: * For each task ID:
* - If no worktree exists / is registered → re-run (fresh worktree later). * - If no worktree exists / is registered → re-run (fresh worktree later).
* - If the worktree working tree is dirty (uncommitted edits) → re-run, * - If the worktree has uncommitted edits to TRACKED files (e.g. an
* preserving the worktree so `createWorktree` reuses it and the agent * interrupted agent mid-edit) → re-run, preserving the worktree so
* continues where it left off. * `createWorktree` reuses it and the agent continues where it left off.
* - If the worktree is clean but has no commits ahead of main → re-run. * Untracked files are ignored here — they never block a merge, and a
* - If the worktree is clean AND has ≥1 commit ahead of main → merge the * worktree whose task work is fully committed is "done" even if it
* branch into main (`--no-ff`), remove the worktree, and report finalized. * carries stray untracked files. Counting `??` entries would strand the
* On merge conflict the merge is aborted (main left clean), the worktree * committed branch in `.ralpi/worktrees/` forever on every resume.
* is preserved, and the task is reported in `conflicts`. * - If the worktree has no commits ahead of main → re-run.
* - If the worktree has ≥1 commit ahead of main → merge the branch into
* main (`--no-ff`) and report finalized. Fully clean worktrees are then
* removed; worktrees that also carry untracked files are kept so that
* (possibly meaningful) uncommitted files aren't destroyed — the next
* fresh-loop sweep cleans them up. On merge conflict the merge is
* aborted (main left clean), the worktree is preserved, and the task is
* reported in `conflicts`.
* *
* This is the self-healing path for an interrupted review-gated loop: * This is the self-healing path for an interrupted review-gated loop:
* tasks that finished (commit + review already saved) but never got their * tasks that finished (commit + review already saved) but never got their
@@ -417,6 +470,7 @@ export function finalizeCommittedWorktrees(
prdKey: string, prdKey: string,
taskIds: string[], taskIds: string[],
): FinalizeResult { ): FinalizeResult {
mainDir = canonicalDir(mainDir);
const result: FinalizeResult = { finalized: [], rerun: [], conflicts: {} }; const result: FinalizeResult = { finalized: [], rerun: [], conflicts: {} };
const mainHead = getGitHead(mainDir); const mainHead = getGitHead(mainDir);
@@ -446,14 +500,15 @@ export function finalizeCommittedWorktrees(
continue; continue;
} }
// Dirty working tree (uncommitted edits, e.g. an interrupted agent) → // Uncommitted edits to TRACKED files (an interrupted agent mid-edit) →
// re-run, keeping the worktree so the agent resumes in place. // re-run, keeping the worktree so the agent resumes in place. Untracked
if (hasUncommittedChanges(wtDir)) { // files alone do NOT count as dirty here (see doc comment above).
if (hasTrackedUncommittedChanges(wtDir)) {
result.rerun.push(taskId); result.rerun.push(taskId);
continue; continue;
} }
// Clean tree but nothing committed ahead of main → nothing to merge. // No commits ahead of main → nothing to merge.
const aheadStr = const aheadStr =
mainHead !== null mainHead !== null
? git(`rev-list --count ${mainHead}..HEAD`, wtDir) ? git(`rev-list --count ${mainHead}..HEAD`, wtDir)
@@ -464,11 +519,17 @@ export function finalizeCommittedWorktrees(
continue; continue;
} }
// Committed + clean → finalize. mergeWorktree aborts on conflict, // Committed + no tracked edits → finalize. mergeWorktree aborts on
// leaving main's working tree clean. // conflict, leaving main's working tree clean.
const merge = mergeWorktree(mainDir, branch); const merge = mergeWorktree(mainDir, branch);
if (merge.success) { if (merge.success) {
// Remove the worktree only when it's fully clean. If it still carries
// untracked files, keep it so that uncommitted work isn't destroyed
// (the branch is merged; the leftover worktree is swept by the next
// fresh-loop cleanup).
if (!hasUncommittedChanges(wtDir)) {
removeWorktree(mainDir, { dir: wtDir, branch, mainDir }); removeWorktree(mainDir, { dir: wtDir, branch, mainDir });
}
result.finalized.push(taskId); result.finalized.push(taskId);
continue; continue;
} }
@@ -480,3 +541,24 @@ export function finalizeCommittedWorktrees(
git("worktree prune", mainDir); git("worktree prune", mainDir);
return result; return result;
} }
/**
* Whether a worktree still holds work worth preserving (committed commits
* ahead of main, or uncommitted changes). Used by the task-failure path so a
* failed/timeout agent's partial output isn't force-deleted with the
* worktree.
*/
export function worktreeHasPreservableWork(
mainDir: string,
wt: WorktreeHandle,
): boolean {
mainDir = canonicalDir(mainDir);
// Any uncommitted changes (tracked edits or untracked files) count — the
// agent may have been mid-write when it failed.
if (hasUncommittedChanges(wt.dir)) return true;
const mainHead = getGitHead(mainDir);
if (!mainHead) return true;
const aheadStr = git(`rev-list --count ${mainHead}..${wt.branch}`, mainDir);
const ahead = aheadStr !== null ? parseInt(aheadStr, 10) : 0;
return !Number.isNaN(ahead) && ahead > 0;
}

View File

@@ -0,0 +1,82 @@
/// <reference types="bun-types" />
import { describe, it, expect, beforeEach } from "bun:test";
import * as fs from "node:fs";
import * as os from "node:os";
import * as path from "node:path";
import { ProgressTracker } from "../src/progress";
/**
* Regression test: two concurrent loops (different PRDs) each run their own
* ProgressTracker. Each instance snapshots the whole state at construction;
* a save() that writes that stale snapshot verbatim would revert the OTHER
* loop's task status changes — tasks wrongly back to "pending" while their
* worktrees carry real work, stranding it on the next resume.
*/
let root: string;
beforeEach(() => {
root = fs.mkdtempSync(path.join(os.tmpdir(), "ralpi-prog-test-"));
});
function prdA(projectDir: string): ProgressTracker {
return new ProgressTracker(
projectDir,
path.join(projectDir, "tasks/a/README.md"),
);
}
function prdB(projectDir: string): ProgressTracker {
return new ProgressTracker(
projectDir,
path.join(projectDir, "tasks/b/README.md"),
);
}
/** Read the on-disk progress state; the file is written by the tracker, so
* a parse failure is a test bug worth surfacing. */
function readState(): Record<string, any> {
const raw = fs.readFileSync(
path.join(root, ".ralpi", "progress.json"),
"utf-8",
);
try {
return JSON.parse(raw) as Record<string, any>;
} catch {
throw new Error(`malformed progress.json:\n${raw.slice(0, 200)}`);
}
}
describe("ProgressTracker multi-PRD save isolation", () => {
it("does not clobber another PRD's task status on save", () => {
const a = prdA(root);
const b = prdB(root);
expect(a.getKey()).not.toBe(b.getKey());
// Loop A marks its task in_progress.
a.markInProgress("01");
expect(a.getTaskStatus("01")).toBe("in_progress");
// Loop B (stale snapshot from before A's update) marks ITS task.
b.markInProgress("02");
// The on-disk state must show BOTH updates.
const raw = readState();
expect(raw.prds[a.getKey()].tasks["01"].status).toBe("in_progress");
expect(raw.prds[b.getKey()].tasks["02"].status).toBe("in_progress");
});
it("preserves other PRD completions when this PRD saves", () => {
const a = prdA(root);
const b = prdB(root);
a.markCompleted("01", 1000);
b.markInProgress("02");
// A completes another task later — A's save must not revert B.
a.markCompleted("03", 500);
const raw = readState();
expect(raw.prds[a.getKey()].tasks["01"].status).toBe("completed");
expect(raw.prds[a.getKey()].tasks["03"].status).toBe("completed");
expect(raw.prds[b.getKey()].tasks["02"].status).toBe("in_progress");
});
});

View File

@@ -0,0 +1,87 @@
/// <reference types="bun-types" />
import { describe, it, expect, beforeEach } from "bun:test";
import * as fs from "node:fs";
import * as os from "node:os";
import * as path from "node:path";
import { ProgressTracker } from "../src/progress";
import { countPRDResumeStats } from "../src/utils";
/**
* Regression test: the resume-selection prompt under-reported task totals
* when multiple loop histories existed. The progress tracker only records
* TOUCHED tasks (started/completed/failed) — never-started tasks are absent
* from prd.tasks, so a naive Object.keys() count missed them entirely, and
* file-checked completions were ignored unless markCompleted had run.
* countPRDResumeStats derives the true total from the parsed PRD file and
* counts checkbox completions too.
*/
let root: string;
beforeEach(() => {
root = fs.mkdtempSync(path.join(os.tmpdir(), "ralpi-stats-test-"));
});
const PRD_CONTENT = `# Test PRD
## Tasks
- [ ] Task one
- [x] Task two
- [ ] Task three
- [ ] Task four
`;
function writePRD(rel: string): string {
const p = path.join(root, rel);
fs.mkdirSync(path.dirname(p), { recursive: true });
fs.writeFileSync(p, PRD_CONTENT, "utf-8");
return p;
}
describe("countPRDResumeStats", () => {
it("reports the full task total from the PRD file, not just touched tasks", () => {
const sourcePath = writePRD("tasks/a/README.md");
const progress = new ProgressTracker(root, sourcePath);
// Simple-checkbox format assigns sequential ids 00-03. Only 00
// (completed) and 02 (failed) were touched by the loop; 01 is checked
// off in the file; 03 was never started.
progress.markCompleted("00", 1000);
progress.markFailed("02", "boom");
const stats = countPRDResumeStats(progress.getState(), sourcePath);
expect(stats.total).toBe(4); // old code reported 2
expect(stats.completed).toBe(2); // 00 via progress + 01 via checkbox
expect(stats.failed).toBe(1);
});
it("does not double-count a task that is both progress-completed and file-checked", () => {
const sourcePath = writePRD("tasks/b/README.md");
const progress = new ProgressTracker(root, sourcePath);
progress.markCompleted("01", 500); // 01 already [x] in the file
const stats = countPRDResumeStats(progress.getState(), sourcePath);
expect(stats.completed).toBe(1);
});
it("falls back to touched-task counts when the PRD file is missing", () => {
const missing = path.join(root, "tasks/gone/README.md");
const progress = new ProgressTracker(root, missing);
progress.markCompleted("01", 1000);
progress.markFailed("02", "nope");
const stats = countPRDResumeStats(progress.getState(), missing);
expect(stats.total).toBe(2);
expect(stats.completed).toBe(1);
expect(stats.failed).toBe(1);
});
it("reports zero for a never-touched PRD with no file", () => {
const missing = path.join(root, "tasks/none/README.md");
const progress = new ProgressTracker(root, missing);
const stats = countPRDResumeStats(progress.getState(), missing);
expect(stats.total).toBe(0);
expect(stats.completed).toBe(0);
expect(stats.failed).toBe(0);
});
});

View File

@@ -0,0 +1,265 @@
/// <reference types="bun-types" />
import { describe, it, expect, beforeEach } from "bun:test";
import { execSync } from "node:child_process";
import * as fs from "node:fs";
import * as os from "node:os";
import * as path from "node:path";
import {
createWorktree,
finalizeCommittedWorktrees,
mergeWorktree,
removeWorktree,
worktreeHasPreservableWork,
} from "../src/worktree";
/**
* Regression tests for worktree resume/finalize behavior:
*
* 1. finalizeCommittedWorktrees merges a committed worktree branch even
* when the worktree carries UNTRACKED files (previously the dirty check
* counted `??` entries, stranding committed code in .ralpi/worktrees/).
* 2. finalize works for tasks that are NOT in_progress (pending) — the
* stranded-work case after an interrupted resume.
* 3. createWorktree reuses an existing worktree under a symlinked project
* path (git porcelain emits realpaths; literal path.join must not be
* compared verbatim).
* 4. worktreeHasPreservableWork keeps failed-task branches alive so a
* timeout doesn't destroy commits the agent already made.
*/
const sh = (cmd: string, cwd: string): string => {
try {
return execSync(cmd, { cwd, encoding: "utf-8" }).trim();
} catch (err) {
throw new Error(
`git cmd failed in ${cwd}: ${cmd}\n${(err as Error).message}`,
);
}
};
const STATE_DIR = ".ralpi";
const PRD_KEY = "prd";
function makeRepo(root: string): void {
sh("git init -q -b master .", root);
sh("git config user.email t@t.co", root);
sh("git config user.name T", root);
sh("echo '# Demo' > README.md", root);
sh("git add -A && git commit -qm init", root);
}
/** Commit work in a worktree and record the commit message. */
function commitInWorktree(wt: { dir: string }, filename: string, msg: string) {
sh(`echo '${filename} content' > ${filename}`, wt.dir);
sh(`git add -A && git commit -qm '${msg}'`, wt.dir);
}
function masterHasFile(root: string, filename: string): boolean {
try {
sh(`git show master:${filename}`, root);
return true;
} catch {
return false;
}
}
let root: string;
beforeEach(() => {
root = fs.mkdtempSync(path.join(os.tmpdir(), "ralpi-wt-test-"));
makeRepo(root);
});
describe("finalizeCommittedWorktrees", () => {
it("merges a committed worktree branch even when untracked files exist", () => {
const wt = createWorktree(
root,
STATE_DIR,
"01",
PRD_KEY,
undefined,
"task one",
)!;
commitInWorktree(wt, "work.txt", "task 01 work");
// The task agent left a scratch file untracked (like build artifacts).
sh("mkdir -p scratch && echo junk > scratch/junk.bin", wt.dir);
expect(sh("git status --porcelain", wt.dir)).toContain("??");
const fin = finalizeCommittedWorktrees(root, STATE_DIR, PRD_KEY, ["01"]);
expect(fin.finalized).toEqual(["01"]);
expect(masterHasFile(root, "work.txt")).toBe(true);
});
it("finalizes tasks that are pending (not in_progress) with committed work", () => {
const wt = createWorktree(
root,
STATE_DIR,
"02",
PRD_KEY,
undefined,
"task two",
)!;
commitInWorktree(wt, "b.txt", "task 02 work");
// Simulate a prior interrupted resume: task reset to pending, branch
// never merged.
const fin = finalizeCommittedWorktrees(root, STATE_DIR, PRD_KEY, ["02"]);
expect(fin.finalized).toEqual(["02"]);
expect(masterHasFile(root, "b.txt")).toBe(true);
});
it("leaves a worktree with uncommitted TRACKED edits for re-run", () => {
const wt = createWorktree(
root,
STATE_DIR,
"03",
PRD_KEY,
undefined,
"task three",
)!;
commitInWorktree(wt, "c.txt", "task 03 work");
// Agent was mid-edit when interrupted: a tracked file modified.
sh("echo more >> README.md", wt.dir);
const fin = finalizeCommittedWorktrees(root, STATE_DIR, PRD_KEY, ["03"]);
expect(fin.finalized).toEqual([]);
expect(fin.rerun).toEqual(["03"]);
expect(masterHasFile(root, "c.txt")).toBe(false);
});
it("does not re-merge an already-merged branch", () => {
const wt = createWorktree(
root,
STATE_DIR,
"04",
PRD_KEY,
undefined,
"task four",
)!;
commitInWorktree(wt, "d.txt", "task 04 work");
expect(mergeWorktree(root, wt.branch).success).toBe(true);
removeWorktree(root, wt);
// Re-create a worktree on the same (now-merged) branch tip: nothing
// ahead of main → re-run, no spurious merge.
const wt2 = createWorktree(
root,
STATE_DIR,
"04",
PRD_KEY,
undefined,
"task four",
)!;
commitInWorktree(wt2, "e.txt", "task 04 more work");
const fin = finalizeCommittedWorktrees(root, STATE_DIR, PRD_KEY, ["04"]);
expect(fin.finalized).toEqual(["04"]);
expect(masterHasFile(root, "e.txt")).toBe(true);
});
it("reports conflicts and preserves the worktree", () => {
const wt = createWorktree(
root,
STATE_DIR,
"05",
PRD_KEY,
undefined,
"task five",
)!;
// Both sides edit f.txt: master AFTER the worktree exists, so the
// branches genuinely diverge and the merge must conflict.
sh(
"echo master > f.txt && git add -A && git commit -qm 'master f.txt'",
root,
);
sh("echo worktree > f.txt", wt.dir);
sh("git add -A && git commit -qm 'task 05 work'", wt.dir);
const fin = finalizeCommittedWorktrees(root, STATE_DIR, PRD_KEY, ["05"]);
expect(fin.finalized).toEqual([]);
expect(fin.conflicts["05"]).toBeTruthy();
// worktree preserved for manual resolution
expect(fs.existsSync(wt.dir)).toBe(true);
});
});
describe("createWorktree resume reuse under symlinked paths", () => {
it("reuses an existing worktree when the project path contains a symlink", () => {
// macOS /tmp → /private/tmp style symlink: git porcelain reports the
// REAL path, path.join keeps the literal one. Reuse must still match.
const realBase = fs.mkdtempSync(path.join(os.tmpdir(), "ralpi-wt-real-"));
const link = path.join(realBase, "link");
fs.mkdirSync(path.join(realBase, "repo"));
fs.symlinkSync(path.join(realBase, "repo"), link);
const symRoot = link;
makeRepo(symRoot);
// Sanity: this is genuinely a symlink situation.
expect(fs.realpathSync(symRoot)).not.toBe(symRoot);
const wt = createWorktree(
symRoot,
STATE_DIR,
"01",
PRD_KEY,
undefined,
"task one",
)!;
commitInWorktree(wt, "a.txt", "task 01 work");
// Resume: createWorktree again must REUSE the registered worktree
// (same dir), not fail and fall through to the main repo.
const reused = createWorktree(
symRoot,
STATE_DIR,
"01",
PRD_KEY,
undefined,
"task one",
)!;
expect(reused.dir).toBe(fs.realpathSync(wt.dir));
const fin = finalizeCommittedWorktrees(symRoot, STATE_DIR, PRD_KEY, ["01"]);
expect(fin.finalized).toEqual(["01"]);
expect(masterHasFile(fs.realpathSync(symRoot), "a.txt")).toBe(true);
});
});
describe("worktreeHasPreservableWork", () => {
it("returns true for a worktree with committed work ahead of main", () => {
const wt = createWorktree(
root,
STATE_DIR,
"01",
PRD_KEY,
undefined,
"task one",
)!;
commitInWorktree(wt, "a.txt", "task 01 work");
expect(worktreeHasPreservableWork(root, wt)).toBe(true);
});
it("returns true for a worktree with uncommitted changes", () => {
const wt = createWorktree(
root,
STATE_DIR,
"01",
PRD_KEY,
undefined,
"task one",
)!;
sh("echo x > junk.txt", wt.dir);
expect(worktreeHasPreservableWork(root, wt)).toBe(true);
});
it("returns false for an empty fresh worktree", () => {
const wt = createWorktree(
root,
STATE_DIR,
"01",
PRD_KEY,
undefined,
"task one",
)!;
expect(worktreeHasPreservableWork(root, wt)).toBe(false);
});
});