PATPengyi Agent Team

Agent Harness Capability Platform

Build agents
that can operate.

PAT turns real tasks into governed Agents with tools, memory, state, evaluation, human authority and production evidence.

Open Learning Research & Development
A modular Agent control center connecting skills, tools, memory, evaluation, security and deployment
Shared infrastructure for every Pengyi Agent Team
Real task firstSchema-bound toolsHuman authorityEvidence before claims

Five independently measurable capabilities

Acquire. Retain. Act. Plan. Improve.

01

RAG

Acquire grounded external knowledge through governed retrieval and context construction.

02

Memory

Retain approved working, episodic, semantic and procedural state.

03

Tools + Skills

Combine atomic interfaces with versioned, tested and permissioned workflows.

04

Planning

Decompose goals, manage state, route actions and stop within explicit limits.

05

Evaluation

Compare task quality, cost, reliability and safety against frozen baselines.

Not another business Agent

The shared operating layer.

Knowledge

Research what Agents can do.

Convert emerging methods into reusable, tested platform capabilities.

Harness

Make execution explicit.

Connect model, context, memory, tools, state and approval as one runtime.

Lab

Experiment before adoption.

Run baselines and controlled experiments before domain teams integrate changes.

Governance

Keep power bounded.

Permissions, budgets, audit logs, human gates, release controls and rollback.

Reusable capability, domain-owned outcomes

One foundation.
Many real-world laboratories.

  1. 01PAT Foundation

    Knowledge, memory, tools, planning, evaluation and shared governance.

  2. 02Domain Teams

    Quant, data, FICC, coding, research and career workflows own business correctness.

  3. 03Real Tasks

    Executed outcomes produce traces, failures and domain-specific benchmark evidence.

  4. 04Adaptation

    Offline experiments propose improvements under regression tests and Human Release Gates.

Prompt is one component, not the system

From requirement to monitored production.

  1. 01Define

    Real task, deliverable, success contract and baseline.

  2. 02Decompose

    Workflow, roles, boundaries, handoffs and failure cases.

  3. 03Connect

    Typed data, tools, APIs, context and memory.

  4. 04Orchestrate

    State machine, bounded retries, escalation and termination.

  5. 05Control

    Permissions, budgets, Human Gates and audit trail.

  6. 06Evaluate

    Frozen tasks, comparable baselines and release thresholds.

  7. 07Operate

    Deployment, monitoring, rollback and failure recovery.

  8. 08Improve

    Feedback, gap analysis and versioned evidence.

Twelve reusable modules

Domain teams inherit the platform, then add their expertise.

01

Architecture

Agent definition and stable boundaries.

02

Skills

Reusable professional workflows.

03

Tools

Schema-bound capabilities and mocks.

04

API Hub

Replaceable provider interfaces.

05

Data

Registries, runs, experiments and evidence.

06

Memory

Working, episodic, semantic and procedural state.

07

Context

Relevant information under explicit budgets.

08

Orchestration

State, routing, retries and handoffs.

09

Evaluation

Quality, cost, reliability and safety evidence.

10

Observability

Traces, logs, metrics and incident context.

11

Security

Least privilege, approval and isolation.

12

Deployment

Release, monitoring, rollback and ownership.

Deterministic control plane

An Agent always knows where it is.

CreatedPlanningValidatingRunningReviewCompleted
L0–L2

Bounded autonomy

Read, analyze, draft and operate in isolated test environments.

L3

Controlled modification

Repository and non-production changes with complete traceability.

L4–L5

Human authority

Publishing, production, capital and trading actions cannot self-approve.

Continuous research, measured promotion

Five Agents improve every Agent Team.

01 / KNOWLEDGE

RAG Agent

Tracks retrieval, Graph RAG, agentic retrieval, context engineering, grounding and citation research.

Papers · repositories · retrieval benchmarks
02 / RETENTION

Memory Agent

Studies working, episodic, semantic and procedural memory, forgetting, storage systems and hardware constraints.

Memory systems · storage · hardware
03 / ACTION

Tool Use Agent

Tracks Tools, Skills, MCP servers, APIs, tool discovery, routing, permissions and failure recovery.

Tools · Skills · MCP · safe execution
04 / EXECUTION

Planning Agent

Studies Agent Harness design, task decomposition, constraints, state machines, loops, termination and recovery.

Harness · planning · orchestration
05 / IMPROVEMENT

Evaluation Agent

Tracks Agent and model benchmarks, validity, contamination, reproducibility and regression evaluation.

Benchmarks · eval adapters · release evidence
Daily research loopScan → Triage → Candidate → Reproduce → Benchmark → Human review → Release → Domain feedback

PAT publishes reusable capability contracts. Domain Teams own business correctness. No Agent can approve its own promotion.