Skip to content

Latest commit

 

History

History
93 lines (72 loc) · 4.48 KB

File metadata and controls

93 lines (72 loc) · 4.48 KB

AI / LLM Integration - PWN::AI

One agent loop, six interchangeable engines. Swap providers by changing one line in ~/.pwn/pwn.yaml; the tool-calling contract is normalized so the agent code never cares which model is behind it.

Multi-provider integration

Supported engines

Engine Client Auth Notes
openai PWN::AI::OpenAI key: function-calling native
anthropic PWN::AI::Anthropic key: tool-use native
grok PWN::AI::Grok key: or oauth: true OAuth = RFC-8628 device-code flow using xAI's public Grok-CLI client id (no secret) - see skill xai_grok_oauth_device_flow
gemini PWN::AI::Gemini key: function-calling native
ollama PWN::AI::Ollama none local - native /api/chat (num_ctx, keep_alive, low-temp + format:'json' on tool turns) and /api/embed for PWN::MemoryIndex
openwebui PWN::AI::OpenWebUI key: (JWT / API token) + base_uri: Open WebUI gateway - OpenAI-compatible /api/v1/chat/completions plus proxied Ollama /ollama/api/* (including embed)

PWN is model-agnostic. ai.<engine>.model is passed straight through to the provider - the codebase and docs deliberately never name a specific model id so you can point each engine at whatever the vendor currently ships (or whatever ollama list shows locally).

Selecting an engine

# ~/.pwn/pwn.yaml
ai:
  active: grok
  grok:
    oauth:
      enroll: true    # first run opens https://accounts.x.ai/... device page
# at runtime
PWN::Env[:ai][:active] = :ollama

Engine-aware behavior

The harness adapts to the class of engine, not the model name:

Concern Frontier (openai · anthropic · grok · gemini) Local / gateway (ollama · openwebui)
PromptBuilder.budget full MEMORY / METRICS / MISTAKES / LEARNING / EXTRO blocks tightened via ai.ollama.prompt_budget (extro off by default)
MEMORY ranking relevance-ranked when a local Ollama embed_model is reachable, else newest-first relevance-ranked via PWN::MemoryIndex (~/.pwn/memory.idx)
Tool schemas shipped all toolsets CORE_TOOLS + top-K keyword matches when ai.agent.tool_router is on
Pre-pass none plan_first numbered tool plan before first dispatch
Intent route always request_intent + LLM/heuristic request_kind (statement | question | autonomous_goal). Short-circuits how-to/questions (text only), greetings/statements (fixed ack), pure recall, and unauthorized recon on all engines; host-evidence Qs (hostname/cwd/whoami) and only true autonomous goals get multi-step TaskSummarizer plans. Critical for ollama/openwebui
Few-shot none Learning.exemplars_for(request) splices a prior successful trace
Dispatch parsing strict tolerant - Levenshtein tool-name repair + JSON5-ish arg cleanup, each repair fingerprinted into Mistakes
Post-answer auto_introspect auto_introspect + fact_check_local_final (auto extro_verify on CVE/version-shaped claims)
Metrics bucket metrics.json[:tools][name][:engines][:<engine>] same - the TOOL EFFECTIVENESS block is per-engine so local telemetry never blends with frontier

Teacher-student reflection

ai.reflect_engine decouples doing from learning-about-doing:

ai:
  active: ollama            # the local model executes tools and answers
  reflect_engine: anthropic # a frontier model writes the durable Memory :lesson

PWN::AI::Agent::Reflect.on temporarily flips the active engine for the introspection call only, so the local model reads back distilled reasoning it could not have produced itself.

Direct client use (no agent)

resp = PWN::AI::Anthropic.chat(
  messages: [{ role: 'user', content: 'Explain CVE-2024-1234 in one line' }]
)
puts resp[:content]

Model diversity in Swarm

Because each persona in agents.yml can override engine:, an agent_debate can pit several different providers against each other - real antagonism, not one model role-playing three voices. The same mechanism backs ai.agent.escalation_persona: when a local model is stuck, Loop.run asks a frontier persona for a 3-line corrective hint and injects it as a synthetic tool result.

See also: pwn-ai Agent · Agent Tool Registry · Swarm · Configuration

← Home