RAG
Acquire grounded external knowledge through governed retrieval and context construction.
Agent Harness Capability Platform
PAT turns real tasks into governed Agents with tools, memory, state, evaluation, human authority and production evidence.
Open Learning Research & Development →
Five independently measurable capabilities
Acquire grounded external knowledge through governed retrieval and context construction.
Retain approved working, episodic, semantic and procedural state.
Combine atomic interfaces with versioned, tested and permissioned workflows.
Decompose goals, manage state, route actions and stop within explicit limits.
Compare task quality, cost, reliability and safety against frozen baselines.
Not another business Agent
Convert emerging methods into reusable, tested platform capabilities.
Connect model, context, memory, tools, state and approval as one runtime.
Run baselines and controlled experiments before domain teams integrate changes.
Permissions, budgets, audit logs, human gates, release controls and rollback.
Reusable capability, domain-owned outcomes
Knowledge, memory, tools, planning, evaluation and shared governance.
Quant, data, FICC, coding, research and career workflows own business correctness.
Executed outcomes produce traces, failures and domain-specific benchmark evidence.
Offline experiments propose improvements under regression tests and Human Release Gates.
Prompt is one component, not the system
Real task, deliverable, success contract and baseline.
Workflow, roles, boundaries, handoffs and failure cases.
Typed data, tools, APIs, context and memory.
State machine, bounded retries, escalation and termination.
Permissions, budgets, Human Gates and audit trail.
Frozen tasks, comparable baselines and release thresholds.
Deployment, monitoring, rollback and failure recovery.
Feedback, gap analysis and versioned evidence.
Twelve reusable modules
Agent definition and stable boundaries.
Reusable professional workflows.
Schema-bound capabilities and mocks.
Replaceable provider interfaces.
Registries, runs, experiments and evidence.
Working, episodic, semantic and procedural state.
Relevant information under explicit budgets.
State, routing, retries and handoffs.
Quality, cost, reliability and safety evidence.
Traces, logs, metrics and incident context.
Least privilege, approval and isolation.
Release, monitoring, rollback and ownership.
Deterministic control plane
Continuous research, measured promotion
Tracks retrieval, Graph RAG, agentic retrieval, context engineering, grounding and citation research.
Papers · repositories · retrieval benchmarksStudies working, episodic, semantic and procedural memory, forgetting, storage systems and hardware constraints.
Memory systems · storage · hardwareTracks Tools, Skills, MCP servers, APIs, tool discovery, routing, permissions and failure recovery.
Tools · Skills · MCP · safe executionStudies Agent Harness design, task decomposition, constraints, state machines, loops, termination and recovery.
Harness · planning · orchestrationTracks Agent and model benchmarks, validity, contamination, reproducibility and regression evaluation.
Benchmarks · eval adapters · release evidencePAT publishes reusable capability contracts. Domain Teams own business correctness. No Agent can approve its own promotion.