Research · Personal AI Benchmark

How do you measure a personal AI?

We built an 11-dimension framework for evaluating personal AI systems across memory, agency, proactivity, reliability, cross-system reasoning and more — then scored Atlas against it, honestly, weaknesses first.

Capability (unweighted)

7.59 / 10

Mean across all 11 dimensions

Availability

1.0 / 10

Reported separately — never blended into capability

Evidence grade

Self-documented

n = 1 · self-instrumented · no third-party audit

Two-axis position

Promising,
not yet gettable

Top-tier capability, near-zero availability

The framework

Eleven dimensions, one honest rubric.

Most "AI assistant" comparisons measure how well something answers a question. A personal AI has to do far more: remember a life, reach into real data, take consequential action, notice things unprompted, and keep working over days. The framework scores each dimension 0–10 with fixed anchors, and — critically — keeps availability on a separate axis, because a capability you can't get is not the same as one you can.

Method. Twelve systems were scored against eleven criteria on a fixed 0–10 rubric — ChatGPT, Claude, Gemini, M365 Copilot, Alexa+, Siri, Perplexity, and the personal-agent cohort (Lindy, Town, Martin, OpenClaw), plus Atlas. Competitor scores are comparative judgments from a market read; Atlas's eleven scores were re-derived directly from its source tree at v2.8.37. Weighted totals use four weighting profiles: Neutral (equal weight) plus three archetypes that emphasize what different users actually value.

Limitations, stated up front. This is n = 1 and self-documented — every Atlas metric is Atlas measuring itself. There is no independent benchmark and no third-party security audit. Scores are comparative judgments, not lab measurements: a 1–3 point swing is directionally meaningful; decimals should not be overinterpreted. This measures capability and reports availability separately — it is not a purchase verdict for any product.

The dimensions

How Atlas scores — the strong and the soft.

Six criteria sit near the ceiling. Five are visibly softer, and each soft number carries the reason it's capped. The honest read matters more than the headline mean.

Memory & context model
Durable, structured knowledge of a life
9.0● hold
Best-in-study — a queryable Life Graph of people, promises and events, now editable from chat and self-healing. Near-max on the rubric.
Personal-data reach
Breadth and depth of read access to real data
9.0● hold
Multi-account mail, calendar and contacts, messages, location, home and vehicle.
Actioning / agency
Executes consequential, real-world side-effects
9.0● hold
Mail / iMessage / WhatsApp send, calendar writes, Resy / OpenTable / barber booking, Vapi calls, Nest and Tesla control. The five newest domains stay correctly notify-only and are not re-scored up.
Proactivity
Unprompted monitoring and useful initiation
9.0● hold
Event Fabric + Executive Brain — now more defensible: the confidence gate is live and the false-positive rate is measured, not asserted.
Cross-system reasoning
Connect evidence across apps, people and time
9.0● hold
Reasoning that spans the whole graph rather than a single app's data.
Multi-step / long-horizon
Plan, wait, retry, recover, finish
8.5● hold
Exactly-once ledger and reply-waits carry long-running work across days; a formerly-orphaned workflow suite now runs as a gated CI job.
Privacy, security & trust
Permissions, injection resistance, credential handling
7.0▲ +1.0
Three of four named debts closed: an errorJson sanitizer across 48 routes, a socket-level SSRF rebind-pin, and timing-safe webhook checks. The injection boundary is partially closed and approval gates stay atomic and fail-closed. Capped: no external audit; the bridge surface is still fragile.
Learning & adaptation
Improvement from corrections and observed preferences
6.5▲ +1.5
Editable memory and graph tools, a post-event "how was it?" preference loop, and a penalty-only calibration flywheel feeding measured dismiss/act/ignore back into the confidence gate. Capped: mostly fixture-exercised; not model-level learning.
Voice & ambient access
Voice naturalness / latency and cross-device presence
6.5▲ +0.5
Voice-approval round-trip fixed — any natural yes/no reliably approves a gated write — plus first-person, direct-address delivery. Capped: no ambient hardware fleet; Vapi phone and in-app realtime only.
Reliability
Evidence-weighted — measured telemetry, evals, uptime, not anecdote
6.0▲ +1.5
A shared error path with exception capture, a dead-man's-switch on 16/16 crons, a 28-scenario scored eval, and p50/p90/p95 latency plus a measured nudge false-positive rate on the dashboard. Capped: all self-reported and owner-gated, no external uptime history, no third-party audit, home-Mac single point of failure stands.
Onboarding & generalization
Can a non-builder get value; does it work for arbitrary users
4.0▲ +2.0
Hardcoded-owner routes collapsed to a centralized owner resolver, plus a setup wizard, install docs, an env-manifest, a bootstrap script, and one real second deployment. Capped hard: onboarding is conceded "unproven" — no cold-start test, owner-specific defaults remain, still single-tenant-per-deployment. This is the field's honest low, not a rounding error.

The field

Twelve systems against the rubric.

Every contestant, all eleven criteria, unweighted mean. Cells are shaded by strength. The shape to notice isn't a single winner — it's that personal AIs cluster differently depending on which columns a person cares about.

Assistant MemReachAgencyProactMultiXSys RelTrustVoiceLearnOnboardMean
Gemini7.598788768.56.59.57.73
Atlas v2.8.37 99998.59 6.07.06.56.54.07.59
ChatGPT776587.576.586.59.57.09
M365 Copilot76.5768887.56.5677.05
Claude77.564.597.577.56697.00
Alexa+6.5877.55.56.55.558.568.56.77
Lindy66.5877766.566.56.56.64
OpenClaw78.5108773.5246.536.05
Town66.56.56.566.55.56367.56.00
Martin5.56.56.56.55.55.554.56.55.575.86
Perplexity5664.56.56.55.53.54.54.58.55.55
Siri (Apple)3.56443.563.5984.585.45
Strong Moderate Weak Atlas's profile is a spike — near-max on the left six columns, softening on the right, lowest on onboarding.

Weighting is the argument

The winner depends on what you weight.

There is no single "best personal AI" — there's a best one for a given set of priorities. Under equal weighting Atlas is #2, behind Gemini. Weight agency, reasoning and long-horizon execution — what an operator or builder actually leans on — and it moves to #1. The honest reading is that weighting drives the ranking, so we show all four.

Neutral

Equal weight, all 11 criteria

  • 1Gemini77.3
  • 2Atlas75.9
  • 3ChatGPT70.9
  • 4M365 Copilot70.5
  • 5Claude70.0
  • 6+Alexa+ · Lindy · OpenClaw · Town · Martin · Perplexity · Siri

A · Ambient consumer

Reliability & Voice weighted heavy

  • 1Gemini76.8
  • 2Atlas75.8
  • 3M365 Copilot70.2
  • 4ChatGPT69.8
  • 5Alexa+68.2
  • 6+Claude · Lindy · Martin · OpenClaw · Town · Siri · Perplexity

B · Executive operator

Agency, reasoning & execution heavy

  • 1Atlas81.8
  • 2Gemini77.2
  • 3M365 Copilot71.8
  • 4Claude70.0
  • 5ChatGPT68.8
  • 6+Lindy · OpenClaw · Alexa+ · Town · Martin · Perplexity · Siri

C · Builder / autonomy

Multi-step & Agency heavy

  • 1Atlas82.2
  • 2Gemini77.0
  • 3M365 Copilot71.2
  • 4Claude69.3
  • 5OpenClaw69.2
  • 6+ChatGPT · Lindy · Alexa+ · Town · Martin · Perplexity · Siri

Across every weighting Atlas lands #1 or #2 — the same rank-stability class as Gemini — while personal-agent peers swing far more. That stability, not any single #1, is the real signal.

The insight

Top-tier capability. Near-zero availability. That gap is the whole point.

The two axes are never blended. On capability, Atlas is at the front of the field. On availability it scores 1.0 / 10 — it can't be bought, and no stranger has yet been shown to self-install to value. "Promising, not yet gettable" isn't a weakness to hide; it's precisely the transition a company is built to close.

Method & limitations

The honest asterisks.

A benchmark is only as credible as the caveats it publishes. These are ours, in full.

⚠ Reported separately — never folded into capability

Availability

1.0 / 10

A setup wizard and one gifted second deployment exist — but it still can't be bought, and no stranger has been shown to self-install to value.

Evidence grade

Self-documented
n = 1, self-instrumented

Every Atlas metric is Atlas measuring itself. No independent benchmark, no third-party security audit.

Two-axis position

Promising,
not yet gettable

Top-tier capability depth; near-zero availability. The work that moves it is availability, not more capability.

  • Held fixed
    Competitors were not re-scored from primary sources. Their numbers are a market read; re-scoring external products without fresh primary evidence would be invention, not measurement. Atlas's #2 neutral rank is relative to that held field.
  • Sensitivity
    Reliability 6.0 is the soft number. The instrumentation is wired but the evidence is self-reported and owner-gated. A strict evidence-weighted skeptic would hold it at 5.0–5.5 — which pulls the Ambient total down but does not flip the Executive or Builder rankings.
  • Not credited
    The five newest domains stay notify-only. Flights, hotels, events, shopping and markets search-and-watch but never transact — so they were not counted as agency. Agency held at 9.0, not raised.
  • Still open
    The onboarding gap is real. It's conceded "unproven." The 2.0 → 4.0 move credits plumbing and one gifted install — not a stranger getting value, which remains the single most load-bearing unmet claim.
  • Scope
    This is a capability benchmark, not a purchase verdict. Decimals should not be overinterpreted; only 1–3 point swings are directionally meaningful.

Personal AI Benchmark · 11 dimensions on a fixed 0–10 rubric across 12 systems · Atlas scored from source at v2.8.37. Weighted totals = Σ(score × weight) ⁄ 10. Capability and availability are measured on separate axes and never blended. Self-documented, n = 1; no independent benchmark or third-party security audit. Scores are comparative judgments, not lab measurements.