The control plane for agentic engineering

Agents do the work. Direct decides how far it goes.

AI agents can already investigate an issue, write the change and open the pull request. What they can't do is tell you whether they should have. Direct is the governed system around them: journeys, gates, evidence, and a line of sight from production back to the backlog.

Coming soon · tracking 55 Cratis repositories today
app.cratis.direct/product/journeys
The Direct Journeys board: every issue as a card in Intake, Investigation, Plan, Implementation or Release, with the things that need a person called out.

Real screenshot from production. The Journeys board with our own backlog on it. Cards that need a person are flagged; cards blocked by reality — like an issue closed mid-journey — say so.

The problem

Speed was never the hard part.

Once agents can write code, the constraint moves. It becomes deciding what they work on, proving what they did, and noticing when the world changed underneath them.

An agent that starts on everything spends tokens on everything.

A label flips to "agent can do this" and work begins. Nobody chose it, nobody capped it, and the bill is the first signal.

With Direct: nothing starts until a person admits it into a journey with a work ceiling. Classification describes an issue; it never starts one.

A merged pull request is not evidence.

"The tests passed" and "the model said it was fine" don't answer why a change was allowed to ship, or under whose rules.

With Direct: every stage ends at a gate that evaluates recorded evidence against versioned rules. The verdict, the evidence and the rule version are kept.

Production talks in a channel nobody has time to read.

The alert fires every five minutes, a human reads it eventually, and an agent could have fixed it in thirty seconds.

With Direct: the alert lands where an agent investigates it at once. If it can't resolve it, a person gets the findings, not a notification.
What Direct is

Five loops. One governed system.

Most tools automate one loop — usually the one that writes code. Engineering is more than that. Direct runs the whole cycle from intent to production and back, and gives each loop the same guardrails: a person decides, the rules are visible, the history is kept.

01

Deliver

ADLC · journeys

The Agentic Development Lifecycle. Issues move from intake through investigation, plan, implementation and release, up to a ceiling you set.

02

Assure

gates · rules · evidence

The assurance engine. Each stage ends at a gate. Rules evaluate evidence. Unknown never means yes.

03

Operate

alerts · investigations

Production signals reach an agent with scoped operational access. Findings come back attached to the alert.

04

Publish

content · outlets

Blog posts, the weekly digest and talks follow their own journey with their own gates, from brief to verified publication.

05

Steer

vision · roadmap · chains

Vision documents ground triage. Roadmap plans are generated from the issues you select. Chains sequence work that depends on a release.

ADLC is the first loop, not the whole product. We call the whole thing agentic operations: governing agents across delivery, production and communication.

01 · Deliver — the journey

Work starts in exactly one way.

A journey is the life of one issue or blog post, recorded as events on its own stream. It moves through stages. Between every pair of stages sits a gate. Only a gate that allows can move it on.

01 / stageIntake
G01
02 / stageInvestigation
G02
03 / stagePlan
G03
04 / stageImplementation
G04
05 / stageRelease
G05
outcomeComplete
Issue journey · 5 gates · nothing on any screen can set a stage directly — every action is a command the journey decides on

Admit, then set a ceiling.

You pick issues from the backlog and choose how far work may proceed: pause after investigation, after the plan, after implementation, or run through to release. The ceiling is where work stops. Every stage up to it still has to pass its gate.

  • Preview before you commit. Direct splits your selection into ready and held, with a reason for each held item.
  • One journey per subject. An issue already in a journey can't be admitted twice.
  • Extend scope later. Raise the ceiling from the board when the gate has earned it.
  • Replies steer the next stage. A comment on the agent's write-up becomes input for the next dispatch instead of starting a second run.
  • Changed intent is judged, not ignored. If the title or body changes after the plan, the Decision Engine says whether the plan still holds. If it can't tell, you decide.
app.cratis.direct/product/journeys?journey=…
A journey opened in the side panel: Intake and Investigation gates passed, Plan in progress, and a drift notice that the issue was closed while its journey is still active.

Drift, caught. This issue was closed on GitHub while its journey was still in Plan. Direct doesn't quietly carry on; it stops the journey and asks whether to abandon or re-open.

Reality check

Journeys stay honest against what actually happened.

Every 15 minutes, and whenever a journey changes, Direct compares the journey's record with the facts of the work. A mismatch becomes drift and stops the journey advancing: subject closed, a pull request merged before Release, a worker that failed without finishing, a publication taken down, or governance that changed under it.

Two journeys today

Issues and blog posts share one engine.

The issue journey runs Intake → Investigation → Plan → Implementation → Release. The blog journey runs Brief → Draft → Editorial review → Publish. Same gates, same board, same designer.

02 · Assure — the assurance engine

A gate decides. Evidence and rules explain why.

A gate collects every rule bound to it, evaluates each against the recorded evidence, maps each finding to an outcome, and takes the most restrictive one. Your organization writes the rules. You can read, simulate and version them.

Five outcomes, ordered by how much they ask of you.

AllowThe work may proceed.
Requires human reviewA person approves, with a reason, for exactly the evidence evaluated.
Requires explanationSomeone explains. The explanation becomes evidence for a new attempt.
Requires remediationThe stage is worked again. After two attempts a person decides.
BlockFix the cause first. A block can't be bypassed by asking for a redo.

Unknown is never a pass. Missing evidence, an evaluator error, low confidence or a result that hasn't arrived all default to human review. A late result for evidence that has since changed is recorded but can never authorize the current work.

app.cratis.direct/product/adlc-designer
The ADLC / Journey Designer showing gate G04, Implementation quality, with its policy and the rules bound to it: Build and tests pass, Static analysis, Intent satisfied, Plan conformance.

The ADLC / Journey Designer. Gate G04 — Implementation quality — with 16 required rules. Deterministic checks, Roslyn analysis and LLM judgement sit in one policy.

RuleEvaluatorWhen it doesn't pass
G01 · Intake
Security screeningDeterministicHuman review
Spam and solicitationLLMBlock
Actionable issueLLMRemediation
G03 · Plan quality
Unplanned scopeLLMExplanation
Breaking changes plannedLLMHuman review
Security controlsDecision EngineBlock
G04 · Implementation quality
Build and tests passDeterministicRemediation
Static analysisRoslynRemediation
Tests not weakenedLLMHuman review
Security scanExternal toolBlock
Review policyCompositeHuman review
G05 · Release
MergedDeterministicBlock

Excerpt of our own active governance — 42 rules and 37 bindings across 5 gates for the issue journey.

Six kinds of evaluator. Pick the cheapest one that can decide.

A rule names one evaluator. Cheap deterministic checks run first in spirit; a language model is asked only where judgement is the point.

Deterministic fields & thresholds Composite all / any / threshold / sequence LLM structured finding + confidence Decision Engine configured choices External tool or job Roslyn build diagnostics
  • A low-confidence verdict counts as Unknown. Only a recognized, explained verdict at or above the configured confidence concludes anything.
  • Tool rules wait. A rule handed to an external scanner stays pending until the result arrives for the evidence evaluated.
  • Not applicable isn't passed. A required gate with no applicable rule holds as a configuration problem.
  • Dismissals are evidence. Dismiss a screening flag as a false positive and the rule passes on re-evaluation, with your decision on file. Screen again and it still blocks? The dismissal stops counting.

Governance is versioned, and only a person activates it.

Draft→ Validate→ Simulate→ Review→ Approve→ Active v2
Journeys keep their version

Activating never rewrites work in flight.

A journey keeps the governance it was admitted under. When a newer version changes a gate it still has to pass, Direct raises drift and lets you adopt it deliberately.

Validate + simulate

Break it on a whiteboard, not in a release.

Validation rejects unreachable stages, rule cycles, bindings with no evidence source and outcome maps that let "unknown" allow. Simulation runs a gate against real or typed-in evidence and never calls a model.

Scopes + provenance

Override in one area without changing everywhere.

Product and Content can override a rule's parameters. Each value shows whether it is inherited, overridden, local or excluded, so you always know where a setting came from.

Classification & intake

Classification describes. Admission decides.

Every issue is classified and screened the moment it is registered. The results are recorded as facts with their provenance. None of them starts work, and none of them blocks you from starting it.

KindBug, task or feature. Changed on the issue itself.
FeasibilityWhether an agent can plausibly do it, or it needs a person.
CriticalityHow much it matters. Feeds priority, not permission.
SemVer impactThe release the change warrants. Shown as a guess, with the question mark.
Model tierWhich capability tier should implement it. Cheap work isn't sent to the expensive model.
LabelsApplied from a known set, weighed by the Decision Engine.
ScreeningSecurity (text written to steer an agent) and content (spam, solicitation, conduct), on every issue and every comment.

A person can always say otherwise.

These are bounded questions with a known set of answers, so Direct doesn't ask an open-ended agent. It asks a decision engine and records the verdict.

  • Your classification sits beside the classifier's. It's recorded as an attributed decision. It doesn't overwrite the machine's verdict.
  • Readiness is separate. Ready, Pending or Held — held only when the issue is closed, screening blocked it, or its body says too little.
  • Unclassified issues can still start. Direct won't make you wait for a label to spend your own budget.
  • Outsiders don't trigger work. An issue from outside the organization is screened, flagged and kept away from agents until you decide.
  • Flagged text is shown as plain text. It's content marked as hostile, so it is never rendered.
app.cratis.direct/product/backlog
The Backlog: issues ordered by priority with repository, intent, classification, readiness and a Start Journey button on every row.

The Backlog. Ordered by priority, dragged to reorder, dropped onto each other to group. Classification and readiness sit beside every row, and Start Journey is the only way work begins.

03 · Operate — closing the loop with production

From an alert to a finding, before anyone opens a laptop.

A watchdog that finds something wrong has two options: wake someone or be ignored. Direct adds a third. The alert goes to an agent that looks at once, resolves what it can, and hands a person the evidence for the rest.

01 · live

Signal

A resource probe, a failed build or a webhook from your systems. Direct reads its own payload and the Discord webhook shape.

02 · live

Deduplicate

One alert per fingerprint. A resource down six hours costs one alert and one investigation.

03 · live

Investigate

A worker with only the operational access you granted: a cluster, a runtime, logs.

04 · live

Resolve or hand over

Resolved by the agent, or Needs attention with findings attached.

05 · direction

Fix through a journey

The diagnosis becomes an issue that enters a governed journey, gates included.

06 · direction

Verify in production

The same resource probe confirms the fix landed, and the alert closes itself.

Live in the productWhere we're taking it

Environments and the things in them.

An operational resource is anything worth knowing the health of: a Kubernetes cluster, a namespace, a database, a Chronicle server. Resources are grouped into environments — production, staging, a customer's installation — and nest, so a namespace is probed through its cluster's credentials.

  • A resource type is a class. Add a technology by writing a class. There's no registry and no list to keep in step, and the Add dialog builds its form from what the type says it needs.
  • Readings are documents. Transitions are events. A probe every minute is 1,440 rows a day that would drown the log. Healthy → unhealthy is a fact and lives forever.
  • Unknown is never reported. A resource added a moment ago hasn't been looked at yet. That isn't an alert.
  • Access is granted, not assumed. With no operational access configured, an agent can only reason from the alert. The product tells you so at the top of the page.
# any system can report — this is the whole integration
POST https://<direct>/webhooks/alerts
{
  "source":      "production",
  "title":       "Loki is crash looping",
  "summary":     "loki-0: CrashLoopBackOff (restarts: 370)",
  "severity":    "critical",
  "fingerprint": "pod:loki-0:CrashLoopBackOff"
}
Operations · active investigations
01
Worker never left the starting stateScheduler · critical · investigating
02
Chronicle server cannot be reachedWatchdog · critical · findings attached
03
Reconciliation sweep has stopped runningWatchdog · needs attention · convert to issue

Illustration of the Operations board layout; rows are simplified.

Watching the watcher

The failures worth a check announce nothing.

A recurring sweep that silently stops looks exactly like "nothing to do". Direct's own watchdog checks for that shape of fault: symptoms that are an absence.

Bring your alerting

Point the webhook you already have.

If your watchdog posts to Discord, pointing it at Direct is the whole integration. No second alerting path to keep working.

Findings, not notifications

A person starts from evidence.

Annotate an alert, resolve it, or turn it into a tracked issue in one step. Browser push covers only what nothing else resolves.

04 · Publish — content for every outlet

What shipped is only half the story.

Releases pile up faster than anyone writes about them. Direct treats a blog post like an issue: it has a brief, a draft, an editorial gate and a publication that gets verified. The weekly digest is assembled from what your repositories actually released.

app.cratis.direct/content/digest
The weekly digest: a drafted summary of the week, extracted themes, and a version-movement table across repositories.

The weekly digest. Themes extracted from the week's releases, a generated description in your voice, and the version movement for every repository. Publish is a button a person presses.

A journey for words, with gates.

01Brief
02Draft
03Editorial
04Publish
  • Brief readiness. A post needs a title and an objective before an agent drafts anything.
  • Draft quality and editorial approval. Same gate machinery as code, with its own rules and its own scope in the designer.
  • Publication verification. Publishing isn't finished until the publication is observed. A taken-down post raises drift.
  • A writer agent, in a worker. Drafts come from the content repository or from the issue that motivated them.
In product

Weekly digest

Collected from releases, analysed for themes, reviewed, published.

In product

Blog

Brief to verified publication with editorial gates.

In product · early

Talks & LinkedIn

Talk material and discovered posts with a suggested comment for each.

Direction

More outlets

Newsletter, release notes, docs and social channels as targets for the same content.

05 · Steer — vision, roadmap, order

Agents follow direction. Somebody has to set it.

A backlog of 800 issues has no strategy. Direct keeps the vision next to the work, so triage and planning are grounded in what you're trying to build.

Vision

Written down, and used.

Associate repositories with a vision. Triage and planning read it, and a vision guard checks changes against it.

Roadmap

Plans generated from the issues you select.

Pick issues, get a plan. Revise it, apply it, or throw it away. The plan is a proposal and it stays one until you apply it.

Chains

Step two waits for proof that step one landed.

A fix goes into a framework; the app that consumes it can only be verified once the package is on the registry. Chains hold the later step until the release exists. The agent is told what it's continuing.

Priority

Drag to rank. Drop to group.

A group is scheduled only when every issue in it is ready. Extra instructions travel with the agent's prompt.

Capacity

Least-burnt account first.

Work draws from a pool of AI accounts within a concurrency bound, and usage is tracked per provider and model.

Steering

Watch a worker. Talk back to it.

A running job streams its console into Direct. Send text into the session to steer it, or read the retained console after it ends.

app.cratis.direct/global
The Direct dashboard at a glance: workers running, awaiting your review, alerts needing attention, builds red, open issues, failed partitions.

At a glance. What's running, what awaits you, what's red. Work that finished stops being clutter.

Under the hood

Tools are not authority.

Direct is not an agent loop with a dashboard bolted on. The trusted application owns state, sequencing, credentials and capacity. The agent is a replaceable, non-root worker holding only what one unit of work needs.

Signals in

Your systems

  • GitHub issues, PRs, builds
  • Webhooks + daily reconciliation
  • Alerts from running systems
  • Resource probes
→
Trusted control plane

Direct

  • Journeys, gates, governance
  • Classification & screening
  • Queue, accounts, limits
  • Event-sourced on Chronicle
→
Bounded workers

Isolated agents

  • Docker or Kubernetes, non-root
  • Own worktree per job
  • Scoped callback token
  • Result back as comment, branch, PR
Credentials

Never in the container spec.

Provider keys, GitHub tokens and operational credentials arrive as a read-only secret file the entrypoint sources and deletes. They don't appear in kubectl get job -o yaml or docker inspect.

Source of truth

GitHub is an integration.

Direct records its own facts first and a reactor translates them into GitHub calls. A failed API call is something to recover, not a reason to lose the fact. Another provider means another adapter.

Your models

Bring your own provider.

Anthropic, OpenAI, Azure OpenAI or any OpenAI-compatible endpoint. Pool several accounts. Each agent runs on one provider or draws from a pool, and never both.

Where Direct sits

Delivery tools ship. Governance tools record. Direct does both, and goes back for production.

Autonomous-delivery products turn a backlog into pull requests. Governance products keep evidence for auditors. Direct's position is that those are the same system: the gate that lets an agent proceed is the evidence a reviewer reads later.

Autonomous delivery toolsGovernance & evidence platformsDirect
Starting pointBacklog to pull requestPipeline events to audit trailIntent to journey, up to a ceiling you set
How work is allowedLabels, approval gates, PR reviewControls checked after the factGates evaluate rules against evidence before the next stage
UncertaintyAgent escalates when unsureCompliance statusUnknown is never a pass, by construction
ProductionRarely in scopeRuntime monitoring for driftAlerts trigger agent investigations with scoped access
Beyond code——Content journeys, digest, roadmap, chains
Rules live inRepo configPolicy engineVersioned governance: draft, simulate, review, activate

A category comparison, not a benchmark. We describe what each type of product sets out to do. We haven't tested competitors head to head.

Roadmap

Where it is, and where it's going.

Direct runs our own engineering today. The "next" and "later" columns are direction, not commitments, and dates move.

Now

running in production
  • Journeys and gates for issues and blog posts
  • ADLC / Journey Designer with versioned governance, validation and simulation
  • Classification and screening on every issue and comment
  • Alerts and resource probes with agent investigation
  • Weekly digest, blog, chains and roadmap plans
  • Worker isolation on Docker and Kubernetes

Next

direction
  • Alert → journey. An investigation's diagnosis becomes a governed issue in one step
  • Environment-aware investigations with per-environment access and approval
  • Governance packs — starter rule sets you can adopt and adapt
  • Evidence export for reviewers and auditors
  • Wider watchdog coverage as the internal checks are re-enabled

Later

exploring
  • Fix verified in production by the same probe that raised the alert
  • Feedback loop into rules. Drift and overrides propose rule changes for review
  • Other issue and code providers through the same provider-neutral commands
  • More content outlets from one brief
  • Customer installations as first-class environments
Honest limits

Where Direct stops.

It's GitHub-first.

Issues, pull requests and identity come from GitHub. The internals are provider-neutral, but no other provider ships yet.

Agents are only as good as the rules you activate.

A gate that has no applicable rule holds. A weak rule set passes weak work, and Direct will show you the verdict either way.

LLM rules are judgement, not proof.

They carry a confidence threshold and never turn an unclear answer into a pass. Where you need proof, bind a deterministic or tool rule.

Production access is opt-in.

Without an operational grant, an alert investigation can only reason from the alert. Some internal watchdog checks are currently paused while we harden them.

It's coming soon.

Direct runs our own work today. Access opens to others in stages, so the login below will only admit accounts we've enabled.

Coming soon

Give agents a brief, a ceiling and a gate.

Already have access? Sign in with GitHub and open your journeys.