
rustyorb skills on Claude Market
2 skills published by rustyorb. Each listing includes a one-command install for Claude Code, OpenClaw, Codex, and Hermes, plus a link to the upstream source.
agent-evaluation
Evaluating LLM agents requires different approaches than traditional software testing because the same input can produce different outputs. It covers behavioral regression tests, capability assessments, and reliability metrics designed to catch issues before production. Even top agents score below 50% on real-world benchmarks.
agent-orchestration-multi-agent-optimize
Coordinates multiple specialized agents to profile, optimize, and monitor multi-agent system performance. Covers context window compression, parallel execution, LLM cost tracking, and latency reduction across database, application, and frontend layers.