I turn vague claims into measurable evidence, design software for the paths that fail,
and contribute focused fixes to emulator projects.
Agent evaluation · resilient interfaces · emulation · developer tooling
Prove your Agent Skill works before you publish it.
A dependency-free CLI and companion Agent Skill for comparing clean baselines, automatic availability, and forced skill use across real models. It measures quality, regressions, tokens, cost, and latency, then produces reviewable JSON, HTML, and repository evidence.
Why it exists
Anyone can publish a SKILL.md and say that it helps. SkillProof asks whether
the result actually improved, whether the improvement survives another model,
what it costs, and whether a later edit caused a regression.
The happy path is a demo. Recovery is the product.
An Agent Skill for building complete user interfaces around loading, empty, offline, stale, partial, conflict, permission, retry, and success states. It helps turn a visually polished screen into something people can actually rely on when reality gets messy.
request → waiting → partial data → failure → recovery → continuity
| Area | What I contribute |
|---|---|
| SharpEmu | Focused C# and Vulkan work around resource lifetime, scheduling boundaries, pipeline validation, render-pass reuse, diagnostics, and regression safety. |
| KytyPS5 | Portable Vulkan tests, CMake and CTest integration, and fixes that behave consistently across GPU drivers. |
| healthcheck | A small CLI for checking repository health, contribution activity, release readiness, and maintainer hygiene. |
| codex-maintainer-skills | Compact workflows that help coding agents contribute to open-source projects with better scope and evidence. |
I prefer changes that are small enough to review, important enough to matter, and backed by a test or a concrete explanation.
Measure the claim. Keep the evidence.
Design the failure. Make recovery obvious.
Change one boundary. Test the real behavior.
- Practical tools over impressive demos.
- Reproducible fixes over speculative rewrites.
- Honest limitations over inflated claims.
- Architecture that makes the next contribution safer.
If one of these projects is useful to you, open an issue, try it on a real project, or tell me where it breaks.
