Research · Personal AI Benchmark

How do you measure a personal AI?

We built an 11-dimension framework for evaluating personal AI systems across memory, agency, proactivity, reliability, cross-system reasoning and more — then scored Atlas against it, honestly, weaknesses first.

Capability (unweighted)

7.91 / 10

Mean across all 11 dimensions

Availability

1.0 / 10

Reported separately — never blended into capability

Evidence grade

Self-documented

n = 1 · self-instrumented · no third-party audit

Two-axis position

Promising,
not yet gettable

Top-tier capability, near-zero availability

The framework

Eleven dimensions, one honest rubric.

Most "AI assistant" comparisons measure how well something answers a question. A personal AI has to do far more: remember a life, reach into real data, take consequential action, notice things unprompted, and keep working over days. The framework scores each dimension 0–10 with fixed anchors, and — critically — keeps availability on a separate axis, because a capability you can't get is not the same as one you can.

Method. Twelve systems were scored against eleven criteria on a fixed 0–10 rubric — ChatGPT, Claude, Gemini, M365 Copilot, Alexa+, Siri, Perplexity, and the personal-agent cohort (Lindy, Town, Martin, OpenClaw), plus Atlas. Competitor scores are comparative judgments from a market read (their dimension scores are unchanged from the prior read); Atlas's eleven scores were re-derived directly from its source tree at v3.2.0 (current build: v3.4.40.2; the scores have not been re-run since). Weighted totals use four disclosed profiles — Neutral (all criteria weight 1); Ambient (Reliability & Voice ×3); Executive (Agency, Cross-system & Multi-step ×3); Builder (Multi-step & Agency ×3) — each a weighted mean ×10, so anyone can reproduce them from the table.

Limitations, stated up front. This is n = 1 and self-documented — every Atlas metric is Atlas measuring itself. There is no independent benchmark and no third-party security audit. Scores are comparative judgments, not lab measurements: a 1–3 point swing is directionally meaningful; decimals should not be overinterpreted. This measures capability and reports availability separately — it is not a purchase verdict for any product.

The dimensions

How Atlas scores — the strong and the soft.

Six criteria sit near the ceiling. Five are visibly softer, and each soft number carries the reason it's capped. The honest read matters more than the headline mean.

Memory & context model
Durable, structured knowledge of a life
9.0● hold
Best-in-study — a queryable Life Graph of people, places, events and promises, editable from chat with provenance and an admin surface; commitments are now a managed lifecycle, not just stored facts. Near-max on the rubric.
Personal-data reach
Breadth and depth of read access to real data
9.0● hold
Multi-account mail, calendar and contacts across every connected identity, messages (now including Slack), a connected note-taker's notes and transcripts, location, home and vehicle.
Actioning / agency
Executes consequential, real-world side-effects
9.0● hold
Mail / iMessage / WhatsApp / Slack send, calendar writes, Resy / OpenTable / barber booking, Vapi calls, Nest and Tesla control. The search-only domains (travel, tickets, shopping, markets) stay correctly notify-only and are not re-scored up.
Proactivity
Unprompted monitoring and useful initiation
9.5▲ +0.5
Event Fabric + Executive Brain, now materially cleaner: notification hygiene (one card, count-change-gated), a managed promise check-in that stays silent until due and self-closes on evidence, and conflict detection that no longer flags the same event mirrored across calendars. The false-positive rate is measured, not asserted.
Cross-system reasoning
Connect evidence across apps, people and time
9.0● hold
Reasoning that spans the whole graph rather than a single app's data.
Multi-step / long-horizon
Plan, wait, retry, recover, finish
9.0▲ +0.5
Exactly-once ledger and reply-waits carry work across days; long-lived follow-up watches read every reply since last contact and extend rather than nag, and the managed promise lifecycle waits days and closes itself. Capped: the messaging bridges still lean on a home Mac.
Privacy, security & trust
Permissions, injection resistance, credential handling
7.0● hold
Three of four named debts closed: an errorJson sanitizer across 48 routes, a socket-level SSRF rebind-pin, and timing-safe webhook checks. The injection boundary is partially closed and approval gates stay atomic and fail-closed. Capped: no external audit; the bridge surface is still fragile.
Learning & adaptation
Improvement from corrections and observed preferences
6.5● hold
Editable memory and graph tools, a post-event "how was it?" preference loop, and a penalty-only calibration flywheel feeding measured dismiss/act/ignore back into the confidence gate. Capped: mostly fixture-exercised; not model-level learning.
Voice & ambient access
Voice naturalness / latency and cross-device presence
6.5● hold
Voice-approval round-trip fixed — any natural yes/no reliably approves a gated write — plus first-person, direct-address delivery. Capped: no ambient hardware fleet; Vapi phone and in-app realtime only.
Reliability
Evidence-weighted — measured telemetry, evals, uptime, not anecdote
6.5▲ +0.5
A shared error path, a dead-man's-switch on 17/17 crons, a scored scenario eval, and p50/p90/p95 latency plus a measured nudge false-positive rate on the dashboard — now joined by a hard guardrail against ever reporting an action it didn't take, and self-closing promises/suggestions that ended the duplicate-card churn. Capped: all self-reported and owner-gated, no external uptime history, no third-party audit, home-Mac single point of failure stands.
Onboarding & generalization
Can a non-builder get value; does it work for arbitrary users
6.0▲ +2.0
The capability-package architecture generalized the shape (connection-gated, not hard-wired to one person), joined by in-app self-service connections and a giftable install — and there are now three real deployments, including a second gifted install and a minimal setup completed by a non-builder in under 15 minutes. Still capped: single-tenant-per-deployment, some owner-specific defaults linger, no public self-serve onboarding, n = 3. Improving fast, but not yet productized.

The field

Twelve systems against the rubric.

Every contestant, all eleven criteria, unweighted mean. Cells are shaded by strength. The shape to notice isn't a single winner — it's that personal AIs cluster differently depending on which columns a person cares about.

Assistant MemReachAgencyProactMultiXSys RelTrustVoiceLearnOnboardMean
Atlas v3.2.0 (current build: v3.4.40.2) 9999.599 6.57.06.56.56.07.91
Gemini7.598788768.56.59.57.73
ChatGPT776587.576.586.59.57.09
M365 Copilot76.5768887.56.5677.05
Claude77.564.597.577.56697.00
Alexa+6.5877.55.56.55.558.568.56.77
Lindy66.5877766.566.56.56.64
OpenClaw78.5108773.5246.536.05
Town66.56.56.566.55.56367.56.00
Martin5.56.56.56.55.55.554.56.55.575.86
Perplexity5664.56.56.55.53.54.54.58.55.55
Siri (Apple)3.56443.563.5984.585.45
Strong Moderate Weak Atlas's profile is a spike — near-max on the left six columns, softening on the right, lowest on onboarding.

Weighting is the argument

The winner depends on what you weight.

There is no single "best personal AI" — there's a best one for a given set of priorities. Under equal weighting Atlas now leads, narrowly; weight agency, reasoning and long-horizon execution — what an operator or builder leans on — and the lead widens. It slips to #2 only under an ambient-consumer weighting that leans on reliability and voice, which is exactly its honest soft spot. The reading is that weighting drives the ranking, so we show all four — with the weights disclosed above.

Neutral

Equal weight, all 11 criteria

  • 1Atlas79.1
  • 2Gemini77.3
  • 3ChatGPT70.9
  • 4M365 Copilot70.5
  • 5Claude70.0
  • 6+Alexa+ · Lindy · OpenClaw · Town · Martin · Perplexity · Siri

A · Ambient consumer

Reliability & Voice weighted ×3

  • 1Gemini77.3
  • 2Atlas75.3
  • 3ChatGPT72.0
  • 4M365 Copilot71.0
  • 5Claude68.7
  • 6+Alexa+ · Lindy · Martin · Town · Siri · OpenClaw · Perplexity

B · Executive operator

Agency, reasoning & execution ×3

  • 1Atlas82.9
  • 2Gemini78.2
  • 3M365 Copilot72.6
  • 4Claude71.8
  • 5ChatGPT71.2
  • 6+Lindy · OpenClaw · Alexa+ · Town · Martin · Perplexity · Siri

C · Builder / autonomy

Multi-step & Agency ×3

  • 1Atlas82.0
  • 2Gemini78.0
  • 3M365 Copilot71.7
  • 4Claude71.3
  • 5ChatGPT70.7
  • 6+Lindy · OpenClaw · Alexa+ · Town · Martin · Perplexity · Siri

Across every weighting Atlas lands #1 or #2 — the same rank-stability class as Gemini — while personal-agent peers swing far more. That stability, not any single #1, is the real signal.

The insight

Top-tier capability. Near-zero availability. That gap is the whole point.

The two axes are never blended. On capability, Atlas is at the front of the field. On availability it scores 1.0 / 10 — it can't be bought, and no stranger has yet been shown to self-install to value. "Promising, not yet gettable" isn't a weakness to hide; it's precisely the transition a company is built to close.

Method & limitations

The honest asterisks.

A benchmark is only as credible as the caveats it publishes. These are ours, in full.

⚠ Reported separately — never folded into capability

Availability

1.0 / 10

A setup wizard, a giftable install, and multiple real deployments exist — but it still can't be bought, and no arbitrary stranger has been shown to self-install to value.

Evidence grade

Self-documented
n = 1, self-instrumented

Every Atlas metric is Atlas measuring itself. No independent benchmark, no third-party security audit.

Two-axis position

Promising,
not yet gettable

Top-tier capability depth; near-zero availability. The work that moves it is availability, not more capability.

  • Held fixed
    Competitors were not re-scored from primary sources. Their numbers are a market read; re-scoring external products without fresh primary evidence would be invention, not measurement. Atlas's neutral #1 is relative to that held field — a narrow lead, not a rout.
  • Sensitivity
    Reliability 6.5 is the soft number. The instrumentation is wired but the evidence is self-reported and owner-gated. A strict evidence-weighted skeptic would hold it at 5.5–6.0 — which pulls the Ambient total down but does not flip the Executive or Builder rankings.
  • Not credited
    The five newest domains stay notify-only. Flights, hotels, events, shopping and markets search-and-watch but never transact — so they were not counted as agency. Agency held at 9.0, not raised.
  • Still open
    The onboarding gap is real. The 4.0 → 6.0 move credits the capability-package architecture and a non-builder's under-15-minute install — real progress, but still not a public, self-serve stranger getting value, which remains the single most load-bearing unmet claim.
  • Scope
    This is a capability benchmark, not a purchase verdict. Decimals should not be overinterpreted; only 1–3 point swings are directionally meaningful.

Personal AI Benchmark · 11 dimensions on a fixed 0–10 rubric across 12 systems · Atlas scored from source at v3.2.0 (current build: v3.4.40.2). Weighted totals are a weighted mean × 10 with disclosed weights. Capability and availability are measured on separate axes and never blended. Self-documented, n = 1; no independent benchmark or third-party security audit. Scores are comparative judgments, not lab measurements.