Skip to content
AOS-RES-006 Research-backed evidence map

Agentic Platforms, Local-First Systems, and Malleable Software

Agentic Platforms, Local-First Systems, and Malleable Software: scope, decisions, requirements, evidence, risks, and traceability for the Agent OS programme.

Agentic Platforms, Local-First Systems, and Malleable Software

This document distinguishes sources, observations, inferences, hypotheses, experiments, and normative decisions so that research cannot silently harden into architecture.

Table of Contents

Purpose and Scope

Area: Research and Evidence.

This document distinguishes sources, observations, inferences, hypotheses, experiments, and normative decisions so that research cannot silently harden into architecture.

This document owns the semantics implied by Agentic Platforms, Local-First Systems, and Malleable Software. It does not assert that every described subsystem already exists. It defines the target model, constraints, evidence needed to trust an implementation, and the boundary with adjacent documents.

Normative Position

  1. Agents operate through typed action providers and explicit capability grants, never ambient UI control.
  2. Trust tiers progress from observation to proposal, reversible execution, confirmed sensitive execution, and bounded autonomy.
  3. Every effect records planner, model, inputs, authority, cost, destination, result, and compensation or irreversibility.

Operating Model

The operating model is contract-first and evidence-driven. A component declares its authority, resources, lifecycle, error model, cancellation and timeout behavior, observability, version, and compatibility promise. Backends are replaceable only when the same conformance suite passes and no forbidden platform type leaks into portable layers.

Implementation proceeds through a reference model or mock, deterministic QEMU evidence where relevant, documentation-first physical hardware, and quality-hardware evidence. Pixel 9 adapters remain quarantined according to ADR-0004.

Requirements

  • R01. Agents operate through typed action providers and explicit capability grants, never ambient UI control.
  • R02. Trust tiers progress from observation to proposal, reversible execution, confirmed sensitive execution, and bounded autonomy.
  • R03. Every effect records planner, model, inputs, authority, cost, destination, result, and compensation or irreversibility.
  • R04. Specify normal, partial, denied, timeout, cancellation, restart, upgrade, and permanent-failure behavior.
  • R05. Expose structured diagnostics without leaking secrets or vendor-specific implementation details.
  • R06. Link material unknowns to a claim and, when testable, an experiment with an owner and gate.
  • R07. Update affected documentation and task data when evidence changes the model.

Failure and Degradation

Degradation must be explicit rather than accidental. The system reports capability absence, reduced quality, unavailable provider, stale data, or unsafe condition through typed states. It must not silently fall back to broader authority, unrestricted legacy execution, unverified firmware, lossy data migration, or irreversible agent action.

Recovery defines what state is retained, reconstructed, re-enrolled, compensated, or intentionally discarded. Unsupported hardware or providers are rejected at binding time where possible.

Evidence and Acceptance

  • Privilege-amplification property tests.
  • Shadow-mode comparison with user behavior.
  • Adversarial destination, budget, confirmation, and rollback tests.
  • Evidence records target identity, hardware revision, firmware, source commit, toolchain, configuration, seed, timestamps, artifacts, expected result, actual result, and reviewer.
  • Acceptance requires the referenced tasks to meet their own criteria; prose completion is not implementation completion.

Implementation Obligations

No current task row references this document directly. Before implementation begins, create an owned task or explicitly mark the document as informational.

Risks and Open Questions

  • Prompt injection can redirect authority.
  • Provider descriptions can misstate external effects.
  • Receipts without enforcement become cosmetic audit logs.
  • Open-question rule: an unanswered high-impact question becomes a claim/experiment record and cannot be hidden in meeting notes.
  • Stop rule: work stops or changes track when legal rights, recovery, debug access, safety, or the required evidence path is unavailable.

Related Documents

Planning Reference Anchors

No additional task-specific anchors are required in this baseline.