Local-capable coding agents. Honest certification boundary.

Your machine. Your models. Your agents. Any screen.

ScreenPilot is the professional Command surface for steering RawrXD local coding agents: workspace-aware, Git-aware, approval-aware, and built around a certified local inference lane.

Cloud/model hosting is a separate future lane. Tonight’s release stays desktop-first.

Command · This PC · local workspace

01

Plan

02

Build

03

Verify

Approval waiting

Network-facing remote control remains fail-closed until EGRESS_001.

Deep2 localGit visibleAgent active
Plan, build, verify, or steer an agent...

Tonight’s drop

Drop-ready means scoped, working, and defensible.

Command-home desktop surface
Local Deep2/GGUF agent lane
Workspace and Git context binding
Live activity, stop, and approval states
Phone/browser control designed, not exposed until egress is certified

Certification posture

Local-first is certified. No-dependency product is not claimed yet.

LOCAL_ONLY_001

PASS / CERTIFIED

Deep2 + GGUF local inference, fail-closed when local policy is violated.

STREAMER_CERT_001

10/10 PASS

TinyLlama Q4_K_M live lane at ~11–13 e2e tok/s.

EGRESS_001

OPEN

Whole-product default-deny egress and residual provider quarantine are not claimed yet.

Internal benchmark snapshot · 2026-08-31

RawrXD ScreenPilot vs top agentic IDEs

Fourteen weighted dimensions across Cursor, Windsurf, GitHub Copilot, Claude, and Zed. The score is a date-stamped internal comparison, with N/A dimensions excluded per product—not an independent industry certification.

Command Home is confirmed live with journal resume, Plan / Build / Agent steering, and Ctrl+Shift+W work mode. Cloud and phone steering remain outside the certified boundary.

RawrXD weighted score

72%

Rank vs five references

#3

Lead dimensions

4

Weighted comparison72 / 100
CursorWindsurfCopilotClaudeZed

What the score means

RawrXD leads where local sovereignty, native execution, and bounded action matter most. The gap is concentrated in cloud agents, codebase retrieval wired into the Command home, marketplace depth, and final UI polish.

Long-tail tools such as Continue, Cline, Cody, and Aider are outside this five-reference scorecard. Product capabilities and plan limits change; re-score before using this as a purchasing or investment claim.

Where RawrXD leads

Sovereignty is the differentiator.

LOCAL_ONLY_001

In-process GGUF / Deep2

Local weights load without an Ollama pull or cloud fetch.

BOUNDARY

Fail-closed command broker

Session bind, write leases, approval queue, and event journal on disk.

WORKFLOW

Two-shell UX

Command-home landing plus Ctrl+Shift+W full Win32 IDE, reversible.

PERFORMANCE

Native Win32 lane

No Electron; optional cloud compile paths remain gated off by default.

Honest gaps vs leaders

The close path is explicit.

The next gains are integration work, not a claim that the open gates are already closed.

GapReferenceClose path
Codebase @ indexing / RAGCursorWire semantic index into the Command-home steer handler
Cloud / phone steeringCursor AgentsREMOTE_CONTROL_001 after EGRESS_001
UI polish / menu E2ECursor / VS CodeP1_UI_MENU_E2E_001; shell phases 5–6
Hosted agent parityWindsurfscreenpilot.tech remains preview; live use needs local :11435
Marketplace DXCopilotWork Mode VSIX; hosted HTML remains honest offline preview mode

Dimension matrix

Evidence, boundary, next move.

The matrix records RawrXD’s current evidence and the comparison note used in the internal scorecard.

DimensionRawrXD evidenceBoundary / comparison note
Agent loop (multi-step autonomy)Broker + stop + journalBounded by LOCAL_ONLY_001 and approvals
Codebase @ indexing / RAGSemantic index in Work ModeNot wired to Command-home steer
Multi-file Composer editsAgent diff in Work ModeCommand-home is chat-first
Local GGUF / in-process inferenceDeep2 + LOCAL_ONLY_001No default Ollama or cloud pull
Egress / air-gap postureFail-closed architectureEGRESS_001 open; quarantine remains explicit
Approval gates + write leasesSession bind + write leaseGit commit/push approval and SharedAgentWorkRegistry
Command-home vs IDE splitPlan / Build / Agent + Ctrl+Shift+WReversible two-shell workflow
Native performance (no Electron)Win32 + MASMZed is the Rust-native reference
Cloud / remote agentsNot exposed yetREMOTE_CONTROL_001 waits for EGRESS_001
Extension marketplaceVSIX in Work ModeHosted HTML is preview-only
Terminal / tool executionPowerShell + bounded agentCapExecute gated in Command mode
Git integration in UIContext bar + railBranch and M/A/D state visible
Pricing / self-hostLocal hardware and self-hostNo cloud IDE subscription
Ship maturity / polishP1 UI ~72%Desktop Preview label stays honest

Next ship

Rebuild with Plan / Build / Agent prefixes and the loopback LocalServer; load GGUF in the hub; smoke on the new SHA; record STATIC_DEP only when SSL is active.

No shortcuts

Cloud lane

Want hosted models later?

Use real GPU compute for model serving, not shared website hosting. Keep the marketing site, app shell, and model API as separate layers so local certification stays clean.

Talk cloud lane