Operating System · $29

MYRIAD — Self-Play RL Engine

Let your agent get better at a task by playing against itself. MYRIAD proposes K angle-varied attempts at a problem, runs them through a model you plug in, scores each with a deterministic verifier (never a self-judge), computes critic-free GRPO group-advantage, and distills the winner into a reusable lesson. It's the self-improvement loop behind modern RL agents, packaged as a clean library you can point at any task with a right/wrong signal. Pure Node, zero dependencies.

MYRIAD — Self-Play RL Engine · console
Self-playactive
Critic-free GRPO grou…active
Deterministic verifieractive
Bring-your-own-model …active
Self-play: K angle-varied
Included
Self-play: K angle-varied
Included
Critic-free GRPO
Included
Deterministic verifier
Included

How it works

The exact pipeline, end to end.

1
Install the system
2
Configure to your stack
3
Run the loop
4
It improves each cycle
Format
Runnable scripts + docs
Delivery
Instant after checkout
License
Yours forever · use on every project
Works with
Claude · ChatGPT · Gemini
In practice · A real-life moment

Self-play: K angle-varied attempts per task, then learn from the spread MYRIAD — Self-Play RL Engine turns that from a chore into a few minutes — and you keep it for every project after.

This vs the alternative

Why it's clearly worth it.

This product
Duct-taped manual workflows
Runs
Unattended — it loops on its own
Only when you babysit it
Improvement
Compounds every cycle
Starts from zero each time
Cost
One-time, yours forever
Stacked monthly SaaS
Ownership
Full source, edit anything
Locked black box

The shift it creates

Without it
  • Duct-taped manual workflows
  • Nothing that compounds over time
  • Rebuilding it from zero each time
With it
  • Self-play: K angle-varied attempts per task, then learn from the spread
  • Critic-free GRPO group-advantage — the modern, stable RL signal
  • Deterministic verifier — scored on real outcomes, never self-graded

Use it to…

Real ways people put MYRIAD — Self-Play RL Engine to work.

Self-play

Self-play: K angle-varied attempts per task, then learn from the spread

Critic-free GRPO group-advantage

Critic-free GRPO group-advantage — the modern, stable RL signal

Deterministic verifier

Deterministic verifier — scored on real outcomes, never self-graded

What's inside

Self-play: K angle-varied attempts per task, then learn from the spread
Critic-free GRPO group-advantage — the modern, stable RL signal
Deterministic verifier — scored on real outcomes, never self-graded
Bring-your-own-model + bring-your-own-verifier
Distills winners into reusable lessons · zero-dep Node · MIT source

Who it's for

Solo founders & creators

ship pro results without a team

Agencies & operators

repeatable for every client

AI builders

drop-in capability

Questions, answered

Instant deliveryYours to ownUse on every projectBuilt to ship results

Ready to put MYRIAD — Self-Play RL Engine to work?

Instant download · one-time price · own forever · all sales final

MYRIAD — Self-Play RL Engine