Research · Personal AI Benchmark
How do you measure a personal AI?
We built an 11-dimension framework for evaluating personal AI systems across memory, agency, proactivity, reliability, cross-system reasoning and more — then scored Atlas against it, honestly, weaknesses first.
Capability (unweighted)
7.59 / 10
Mean across all 11 dimensions
Availability
1.0 / 10
Reported separately — never blended into capability
Evidence grade
Self-documented
n = 1 · self-instrumented · no third-party audit
Two-axis position
Promising,
not yet gettable
Top-tier capability, near-zero availability
The framework
Eleven dimensions, one honest rubric.
Most "AI assistant" comparisons measure how well something answers a question. A personal AI has to do far more: remember a life, reach into real data, take consequential action, notice things unprompted, and keep working over days. The framework scores each dimension 0–10 with fixed anchors, and — critically — keeps availability on a separate axis, because a capability you can't get is not the same as one you can.
Method. Twelve systems were scored against eleven criteria on a fixed 0–10 rubric — ChatGPT, Claude, Gemini, M365 Copilot, Alexa+, Siri, Perplexity, and the personal-agent cohort (Lindy, Town, Martin, OpenClaw), plus Atlas. Competitor scores are comparative judgments from a market read; Atlas's eleven scores were re-derived directly from its source tree at v2.8.37. Weighted totals use four weighting profiles: Neutral (equal weight) plus three archetypes that emphasize what different users actually value.
Limitations, stated up front. This is n = 1 and self-documented — every Atlas metric is Atlas measuring itself. There is no independent benchmark and no third-party security audit. Scores are comparative judgments, not lab measurements: a 1–3 point swing is directionally meaningful; decimals should not be overinterpreted. This measures capability and reports availability separately — it is not a purchase verdict for any product.
The dimensions
How Atlas scores — the strong and the soft.
Six criteria sit near the ceiling. Five are visibly softer, and each soft number carries the reason it's capped. The honest read matters more than the headline mean.
errorJson sanitizer across 48 routes, a socket-level SSRF rebind-pin, and timing-safe webhook checks. The injection boundary is partially closed and approval gates stay atomic and fail-closed. Capped: no external audit; the bridge surface is still fragile.The field
Twelve systems against the rubric.
Every contestant, all eleven criteria, unweighted mean. Cells are shaded by strength. The shape to notice isn't a single winner — it's that personal AIs cluster differently depending on which columns a person cares about.
| Assistant | Mem | Reach | Agency | Proact | Multi | XSys | Rel | Trust | Voice | Learn | Onboard | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini | 7.5 | 9 | 8 | 7 | 8 | 8 | 7 | 6 | 8.5 | 6.5 | 9.5 | 7.73 |
| Atlas v2.8.37 | 9 | 9 | 9 | 9 | 8.5 | 9 | 6.0 | 7.0 | 6.5 | 6.5 | 4.0 | 7.59 |
| ChatGPT | 7 | 7 | 6 | 5 | 8 | 7.5 | 7 | 6.5 | 8 | 6.5 | 9.5 | 7.09 |
| M365 Copilot | 7 | 6.5 | 7 | 6 | 8 | 8 | 8 | 7.5 | 6.5 | 6 | 7 | 7.05 |
| Claude | 7 | 7.5 | 6 | 4.5 | 9 | 7.5 | 7 | 7.5 | 6 | 6 | 9 | 7.00 |
| Alexa+ | 6.5 | 8 | 7 | 7.5 | 5.5 | 6.5 | 5.5 | 5 | 8.5 | 6 | 8.5 | 6.77 |
| Lindy | 6 | 6.5 | 8 | 7 | 7 | 7 | 6 | 6.5 | 6 | 6.5 | 6.5 | 6.64 |
| OpenClaw | 7 | 8.5 | 10 | 8 | 7 | 7 | 3.5 | 2 | 4 | 6.5 | 3 | 6.05 |
| Town | 6 | 6.5 | 6.5 | 6.5 | 6 | 6.5 | 5.5 | 6 | 3 | 6 | 7.5 | 6.00 |
| Martin | 5.5 | 6.5 | 6.5 | 6.5 | 5.5 | 5.5 | 5 | 4.5 | 6.5 | 5.5 | 7 | 5.86 |
| Perplexity | 5 | 6 | 6 | 4.5 | 6.5 | 6.5 | 5.5 | 3.5 | 4.5 | 4.5 | 8.5 | 5.55 |
| Siri (Apple) | 3.5 | 6 | 4 | 4 | 3.5 | 6 | 3.5 | 9 | 8 | 4.5 | 8 | 5.45 |
Weighting is the argument
The winner depends on what you weight.
There is no single "best personal AI" — there's a best one for a given set of priorities. Under equal weighting Atlas is #2, behind Gemini. Weight agency, reasoning and long-horizon execution — what an operator or builder actually leans on — and it moves to #1. The honest reading is that weighting drives the ranking, so we show all four.
Neutral
Equal weight, all 11 criteria
- 1Gemini77.3
- 2Atlas75.9
- 3ChatGPT70.9
- 4M365 Copilot70.5
- 5Claude70.0
- 6+Alexa+ · Lindy · OpenClaw · Town · Martin · Perplexity · Siri
A · Ambient consumer
Reliability & Voice weighted heavy
- 1Gemini76.8
- 2Atlas75.8
- 3M365 Copilot70.2
- 4ChatGPT69.8
- 5Alexa+68.2
- 6+Claude · Lindy · Martin · OpenClaw · Town · Siri · Perplexity
B · Executive operator
Agency, reasoning & execution heavy
- 1Atlas81.8
- 2Gemini77.2
- 3M365 Copilot71.8
- 4Claude70.0
- 5ChatGPT68.8
- 6+Lindy · OpenClaw · Alexa+ · Town · Martin · Perplexity · Siri
C · Builder / autonomy
Multi-step & Agency heavy
- 1Atlas82.2
- 2Gemini77.0
- 3M365 Copilot71.2
- 4Claude69.3
- 5OpenClaw69.2
- 6+ChatGPT · Lindy · Alexa+ · Town · Martin · Perplexity · Siri
Across every weighting Atlas lands #1 or #2 — the same rank-stability class as Gemini — while personal-agent peers swing far more. That stability, not any single #1, is the real signal.
The insight
Top-tier capability. Near-zero availability. That gap is the whole point.
The two axes are never blended. On capability, Atlas is at the front of the field. On availability it scores 1.0 / 10 — it can't be bought, and no stranger has yet been shown to self-install to value. "Promising, not yet gettable" isn't a weakness to hide; it's precisely the transition a company is built to close.
Method & limitations
The honest asterisks.
A benchmark is only as credible as the caveats it publishes. These are ours, in full.
⚠ Reported separately — never folded into capability
Availability
1.0 / 10
A setup wizard and one gifted second deployment exist — but it still can't be bought, and no stranger has been shown to self-install to value.
Evidence grade
Self-documented
n = 1, self-instrumented
Every Atlas metric is Atlas measuring itself. No independent benchmark, no third-party security audit.
Two-axis position
Promising,
not yet gettable
Top-tier capability depth; near-zero availability. The work that moves it is availability, not more capability.
- Held fixedCompetitors were not re-scored from primary sources. Their numbers are a market read; re-scoring external products without fresh primary evidence would be invention, not measurement. Atlas's #2 neutral rank is relative to that held field.
- SensitivityReliability 6.0 is the soft number. The instrumentation is wired but the evidence is self-reported and owner-gated. A strict evidence-weighted skeptic would hold it at 5.0–5.5 — which pulls the Ambient total down but does not flip the Executive or Builder rankings.
- Not creditedThe five newest domains stay notify-only. Flights, hotels, events, shopping and markets search-and-watch but never transact — so they were not counted as agency. Agency held at 9.0, not raised.
- Still openThe onboarding gap is real. It's conceded "unproven." The 2.0 → 4.0 move credits plumbing and one gifted install — not a stranger getting value, which remains the single most load-bearing unmet claim.
- ScopeThis is a capability benchmark, not a purchase verdict. Decimals should not be overinterpreted; only 1–3 point swings are directionally meaningful.
Personal AI Benchmark · 11 dimensions on a fixed 0–10 rubric across 12 systems · Atlas scored from source at v2.8.37. Weighted totals = Σ(score × weight) ⁄ 10. Capability and availability are measured on separate axes and never blended. Self-documented, n = 1; no independent benchmark or third-party security audit. Scores are comparative judgments, not lab measurements.