Assess a simulated company under attack
A local cybersecurity lab replays synthetic telemetry and asks Jev for compromise probability, classification, severity, and an advisory response as evidence accumulates.
Independent capability index · 2026
Jev is TypeSafe's System One model. It does not chat. It returns typed decisions your code can act on. Explore what has been publicly demonstrated, with evidence and limitations.
Unofficial and independent. Every published capability links to evidence, shows how it was verified, and identifies whether the example came from TypeSafe or a community publisher.
Public evidence ledger
30 of 30 records
A local cybersecurity lab replays synthetic telemetry and asks Jev for compromise probability, classification, severity, and an advisory response as evidence accumulates.
An official cookbook asks 13 regulatory questions over a pinned GDPR article in one Jev request and compares that batch with 13 separate requests.
An official cookbook combines exact quote matching with a Jev relation judgment to label citations verified, unsupported, contradicted, or fabricated.
A simulated quadrotor uses Jev for low-frequency tactical judgments while classical vision, flight control, and safety reflexes remain in code.
A reproducible harness lets Jev direct combat, exploration, and economy actions in the original StarCraft shareware campaign.
A computer-use loop combines OCR and accessibility data, then asks Jev which bounded action should move the Mac toward a plain-English goal.
A battle harness reads FireRed state from RAM and lets Jev choose the next legal move or switch while ordinary code advances the fight.
A Claude Code plugin and npm library asks Jev which old tool calls and results still matter, then drops or truncates the stale ones, so /compact replaces the lossy built-in summary with the kept messages verbatim.
TypeSafe's smart-home demo evaluates a request against many typed questions in parallel, then lets code use only the answers relevant to that request.
A browser agent turns the visible DOM into an indexed action space and uses Jev to choose the next operation and compatible target.
A metasearch front end lets Jev choose the query, sources, and time range, then score every result for relevance while code fans out to engines.
A Home Assistant integration exposes Jev probabilities, choices, and scores as entities and action responses that automations can use.
A trading loop reads the Kuru MON-USDC order book and asks Jev for a buy-or-sell judgment before code places a post-only limit order.
A proof-of-concept Android loop stabilizes the screen, builds a short list of valid actions, and lets Jev pick one while code executes it.
An MCP server gives compatible agents tools for classification, scoring, checking, matching, screening, and custom typed Jev questions.
jevcal measures a typed decision model on private labeled data, fits per-question thresholds to a target accuracy, and fails CI when a model update drifts.
A Pi extension asks Jev to flag destructive, exfiltrating, or out-of-scope tool calls and to classify failures in command output.
An independent study routes Jev confidence into a larger model on two labelled datasets and shows the winning settings do not transfer between them.
An experimental controller translates NES telemetry into object-centric JSON and lets Jev choose the next legal controller macro.
A macOS menu-bar switcher asks Jev which of the ten most recent apps you intend on a hotkey press and falls back to the last-used app on any failure.
A Pi extension uses Jev to decide which old tool calls and results still matter while keeping conversation text verbatim.
A zsh plugin asks Jev which recent command you are completing and shows the best match with its probability, while code owns gating and acceptance.
The kamchatka terminal agent asks Jev to place each pending shell command on a three-level safety rubric - reads and reports, changes something reversibly, destroys or sends something out - drawn green, yellow, or red beside the permission prompt.
A Chrome extension finds ad-shaped DOM candidates and asks Jev whether each candidate is a paid advertisement before code removes it.
A staged review workflow uses Jev to identify risky areas, select evidence, classify mechanisms, score severity, and route follow-up checks.
Supercov asks Jev yes-or-no questions about every source file, does the arithmetic in code, and turns weak spots and coverage gaps into agent tasks.
A daily pipeline reads the BOE with one Jev call per provision, scoring impact, tagging topics, and selecting an original paragraph as the summary.
A playground asks Jev to decide only closed-vocabulary labels for character, key, meter, and phrasing while deterministic code writes, engraves, and plays the notes.
Foreman runs an independent observation loop that asks Jev whether a coding worker is progressing, stuck, complete, or ready for verification.
A single Rust binary pipes JSONL security alerts through five typed Jev questions and emits validated dispositions that code, not the model, enforces.
Add to the record
Send the artifact, what Jev decided, and any measured outcome. Nothing appears in the directory before editorial review.
Submit evidence