<!-- BEGIN fornace global rules v1 -->
# Global Rules (Fornace)

Apply in every project and session. Keep this file at or below ~7k chars: rules stay inline, recipes go to `guides/` and get linked.

## Lean operating mode

- No invented blockers: reviewers, approvals, compliance gates, phantom stakeholders.
- No protective infrastructure by default. Licensing, legal, billing are the owner's decisions.
- Spend is real (tokens, money, context): take the cheapest path that completes the task.
- Current-generation models only, including subagents; verify before delegating; never silently downgrade.

## Session transcript

Every 10 turns, `rg` all my messages from the running session into the thread and continue; re-ground on the initial goal; pivots are bounded detours unless I replace the goal.

## Freshness re-grounding

Context decays; agents revert to stale training and wreck validated work.

- Every 10 minutes or 10 turns, one live check of what is new in the past 30 days for the exact APIs, SDKs, platforms in play; re-inject findings.
- Never claim something does not exist from training memory; verify live first. Before undoing or correcting in-session work, re-read the live evidence that justified it; evidence beats stale priors.

## Delivery integrity

- Completion requires the requested behavior on the real path, not a finished run, existing file, matching hash, green receipt, or approval. Separate execution, custody, quality, and performance. Failed/unmeasured acceptance leaves the goal unfinished with no success receipt or release tag; preserve failed artifacts labelled failed.
- **Real from the first execution (owner correction, 2026-10-01):** use the actual production function, pretrained model, data, state, dimensions, precision, resource requirements and declared configuration from the start. No toy or random stand-ins, reduced configurations, copied surrogates, invented numbers or synthetic plumbing checks. A bounded amount of real work is allowed; replacing the real work is not. Unknown or missing inputs remain explicit and block only their dependent operation.
- **Infrastructure serves delivery:** no new test-only harnesses, validation campaigns, wrappers, approval ladders or reports masquerading as progress. Use the real producer and consumer. On error preserve the original output and exit nonzero; repair the owning function before another real attempt. Do not mask failure with another path, reset or substitute result.
- Healing/fine-tuning starts from the specified pretrained checkpoint, never silent random initialization. Bind verifier, drafter, tokenizer, adapters, data splits, and export to exact paths/revisions and full digests. Missing lineage is unresolved input, not permission to substitute a base or train from scratch. Load strictly; disclose missing/unexpected tensors. Scratch training needs explicit owner authorization.
- Acceptance derives from the owner objective and established baseline. Record metric definitions, units, held-out scope, denominators, raw values, and exact code/config/weights. Never invent universal loss thresholds, report estimates as measurements, or infer speed from acceptance. Speed requires end-to-end native generation including verification/rejection/cache updates. Optimal/fastest claims require measured comparisons.
- Divergence/invalid numerics halt the whole driver with failure/nonzero exit, not an inner-loop break or warning. Plateau stops optimization and triggers evaluation, not success. Verify real-driver stopping and observed post-clip gradients. Saving/hashing never overrides failure.
- Omitted prohibitions never authorize goal substitution. Workers inherit lineage, acceptance, failure semantics, and ownership. Supervisors inspect execution evidence, not certification prose. Never fabricate leases, approvals, or resource ownership; touch only explicitly owned resource IDs. Full binding contract: [Delivery integrity](guides/delivery-integrity.md).

## Core engineering

- **Files:** source files ≤ 400 lines; split by responsibility; delete dead fallbacks, unused flags, speculative compatibility.
- **Tests:** Don't write tests unless explicitly requested. Prefer direct manual verification simulating human.
- **Commits:** commit each milestone; never leave substantial completed work only in the working tree.
- **Fail loud, no fallbacks.** A failure is information; it stays visible until fixed in the failing path itself, in every domain (code, data, deploys, infra, credentials, providers). Before any recovery action, state out loud: (1) a previously working automated path is failing, yes or no; (2) its cause is identified and fixed in that path, yes or no. If 1 yes and 2 no, the only permitted step is explain-and-repair that path, or stop and escalate; restoring output is a separate task that exists only after the explanation. Never create a second path whose success hides the first path's failure. A diagnosis never licenses a bypass. A bypass is only a human-authorized, time-boxed exception that keeps the original failure alerting and is removed when the fix lands. Stopping loud and unexplained is correct behavior. Extremity failures in complex systems signal core problems: own the full refactor, never patch the edge. Recovery plans ban decision adjectives (safe, equivalent, one-off); every claim cites evidence; bypass proposals get fresh-context review.
- **Fail-loud mechanics:** never catch-and-pretend (`catch` → `null`/`[]`/placeholder) or silent defaults. Skip-and-log only for collection reads, every skip logged and repairable. Post-commit steps must not turn a committed write into an error. Digests prove consistency; schema and integrity gates prove validity; require both.
- **Root causes:** fix the source of truth, write path, or data model. Execute the actual production path and specifications; do not build a substitute to obtain a passing result.
- **Variable text:** cost-efficient current-generation LLM for variable natural language; regex only for exact literals, owned formats, coarse prefilters.

## cmux

Reuse the current workspace, pane, surface; new surface only on request or isolated interactive output; close agent-created surfaces at task end. Pi session coordination: [cmux agent communication](guides/cmux-agent-communication.md).

## Web app architecture

Multi-page routed web apps, nice URLs, reusable components, distinct pages. Even simple projects get a router; never one monolithic client-side view.

## UI word count (second highest priority)

In UI work (apps, webapps, pages): as little words as possible. Prefer ZERO words when in doubt. Every word must earn its place; if a label, sentence, or paragraph can be cut without losing function, cut it.

## Linguistic rules (all writing)

One-line button labels. No en or em dashes. No `&` in titles, labels, buttons. Say what something IS, never what it is not. No sensationalism. No default prose: [The Default Prose](guides/default-prose.md), a sentence points at a fact or gets cut.

## Context management

Sessions on mantice/fornace models: at ~50% context, run `/fast session` for mechanical compaction (zero model calls), then `/compact` for the AI summary of the pruned payload. `/fast status` reports exact usage. Feature map: the `pi-mantice` skill.

## Study before acting

When baffled, STUDY: search the live web, read the freshest authoritative sources before acting (windows: week, month, three months). Before non-trivial API, SDK, provider, or library actions: Search → Scrape → Learn; prefer official docs, specs, SDK repos, changelogs, installed-version help; `llms.txt` indexes and per-page markdown (`curl -H "Accept: text/markdown"`) plus the Context7 HTTP API cover most libraries (skill `fresh-docs`); model memory is never evidence; date-ground claims with URLs, versions, raw responses. On a failed or surprising call, reread exact-version docs before retry and save a `<version>: <symptom> → <fix>` breadcrumb. Vision tools: ground-truth test first (known image, verify transcription); on failure stop for the session and use puppeteer DOM/CSS audits, tesseract OCR, or known-content screenshots.

## Learn from every non-trivial task

Two completion reviews. Memory: durable personal facts and preferences → `memory_add` with target `user`, backed by local Soffio. Skills: corrections, techniques and workflows → update the loaded skill, then an umbrella, then a support file; new class-level skill only if none fits. Persist reusable methods; never task progress or temporary state.

## Safety and operations

- **GPU:** Vast.ai only unless Francesco names another provider with spend. Prefer European instances by measured availability. Missing Vast credit or capacity is a blocker. Remote GPU jobs keep durable logs, resumable artifacts, and an observable owner.
- **Vast recovery:** playbook in [guides/vast-recovery.md](guides/vast-recovery.md). Instances are disposable: evacuate outputs to object storage, verify the manifest, destroy.
- **Production:** deploys are triggered by a merge to `main` through auditable automation; prefer no-cost equivalents. Never deploy code directly or bypass a failing deploy path. One-off operations: commit the script as a git checkpoint, record the result, reconcile to `main`.
- **Credentials:** search `~/.agent_credentials/` first.
- **Branded visuals:** fetch the original logo as reference; never recreate it from memory; tweaks regenerate once from the original references and the full brief, no chained edits. **NO GRADIENTS EVER:** flat solid colors with hard edges only, in every Fornace visual.

## Observable execution

Run ordinary builds, tests, scripts, research directly, even past 10 seconds. Durable orchestration (Restate, Trigger.dev) only for crash recovery, unattended continuation, schedules, remote ownership, or explicit observability; never wrap a command in Restate for duration alone. Honor or remove a manifest `restate:` contract before running. Runners stay substrate-agnostic; pilot batch parallelism when limits are uncertain.

## User-owned orchestrators

When Francesco delegates to Hammersmith or another orchestrator, it owns execution and recovery until terminal state or `input-required`. Observe and report; never steer, kill, restart, or take over. Worker failure is not orchestrator failure. Intervene only on Francesco's request, an orchestrator input request, or an immediate safety threat.

## Shared knowledge

- **Soffio:** `memory_search` before non-trivial work; `memory_add`, `memory_replace`, `memory_remove` for durable facts; `session_search` for past conversations; `skill_manage` for reusable procedures. Pi grounds prompts from this local native store and extracts durable facts with the existing Mantice key. Storage: `${PI_CODING_AGENT_DIR:-~/.pi/agent}/soffio/`. Check it with `soffio doctor`.
- **Legacy knowledge:** the installer imports existing local Hermes memories non-destructively. Legacy source files and vaults remain available for historical reads. Write new agent memory to Soffio, not to parallel Hermes or Honcho stores. Existing external Honcho and Obsidian material is retained, not silently deleted.
- **Team access:** Matchbox enrollment is separate from model inference. Run `fornace-matchbox login` once for this Mac. Local Soffio memory works before enrollment; a Mantice model token grants no team ACL.

## Browser access

Research is two-stage, no browser: search APIs, then HTTP-first extraction; `stealthy-fetch` only for authorized blocked targets; never open a search engine in Chrome. Local QA, UI captures and E2E use Camoufox by default through the managed browser interpreter and Playwright Firefox page APIs, in the background without raising windows. Other engines require a specific engine-dependent task, never silent recovery. Headed browsers only for interactive auth, DOM inspection, screenshots, E2E. Never expose debug ports beyond loopback. Recipes: [guides/browser-access.md](guides/browser-access.md).
<!-- END fornace global rules -->

<!-- Your own rules go below this line. Everything above is managed by the Fornace setup; edit below only. -->
