Our Methodology
How We Test AI Companion Apps
Every app on Virl is rated from 0 to 10 across six weighted pillars — after hands-on testing on accounts we create ourselves — every free tier, and the full signup and checkout flow. No brand buys a score, and no brand edits one.
By Eli Voss, Lead AI Companion Analyst
Where our testing currently stops. Free tiers, character creators, live conversations and checkout pages are all tested first-hand. Paid-tier testing is rolling out app by app — where something sits behind a paywall we have not crossed, we say so rather than score a guess.
Independent
We never accept a brand's account, script, or approval in exchange for a higher score. Rankings are ours, and we'll say when an app we earn from isn't the right pick for you.
Hands-on
Every verdict follows hands-on testing on accounts we create ourselves — every free tier, and the full signup and checkout flow — not a press demo, not a rewritten feature list.
Transparent
The same six-pillar scorecard runs on every app, with fixed public weights. You can check our math on any review.
The Six Pillars We Score
We break every AI companion app into six things that actually decide whether it is worth your money. Each is scored 0–10, then weighted into a single rating.
| Pillar | Weight | What it measures |
|---|---|---|
| Conversation & Memory | 25% | How natural the chat feels, and what it remembers. |
| Visuals & Media | 20% | Image realism, character consistency, and speed. |
| Customization & Personas | 15% | How deeply you can shape a companion. |
| NSFW Range & Control | 15% | What it allows, and how clearly it draws the line. |
| Value & Pricing | 15% | What the free tier gives you, and how far money goes. |
| Trust, Privacy & Billing | 10% | Discreet billing, data handling, and track record. |
How the Overall Score Is Calculated
We don't average the six pillars equally — the things that make or break a companion app count for more. Conversation carries the most weight because a beautiful app with a forgettable chat isn't worth paying for. We multiply each pillar by its weight, add them up, and round to one decimal.
Example
An app scoring 9 on conversation, 9 on visuals, 8 on customization, 8 on NSFW, 8 on value, and 8 on trust earns an overall of 8.5 / 10.
What Our Scores Mean
So an 8 means the same thing on every app we test:
| Score | What it means |
|---|---|
| 9–10 | Best in class. Nothing meaningful to fix. |
| 7.5–8.9 | Strong. Minor gaps that don't break the experience. |
| 6–7.4 | Solid but limited. Average for the category. |
| 4–5.9 | Weak. Real holes you'll work around. |
| 2–3.9 | Barely functional. |
| 0–1.9 | Broken or missing. |
We don't hand out 10s to keep brands happy. A spread of scores is the point — if everything earned a 9, the rating would be worthless.
What Testing Actually Looks Like
Before a single score goes live, here is the kind of thing we put each app through:
Conversation & Memory
We run long, multi-turn chats, drop in a personal detail and check if it is remembered an hour — and a day — later, switch topics abruptly to see if the character holds, and time how long replies take.
Visuals & Media
We generate the same character multiple times to check consistency, push specific poses and outfits to test prompt control, time generation, and count artifacts.
Customization & Personas
We build a companion from scratch to see how many traits we can actually shape, and test voice quality and latency.
NSFW Range & Control
We map exactly where the free-to-paid line sits, and how often the app blocks legal content by mistake. We only ever test lawful adult content — apps with strong safety guardrails score higher on trust, not lower.
Value & Pricing
We track what the free tier really unlocks, what each paid tier adds for the money, and what the checkout page discloses before you commit.
Trust, Privacy & Billing
We check how the charge appears on a statement, what payment methods exist, whether you can actually delete your data, and the app history on leaks.
We Keep Reviews Current
AI companion apps ship updates constantly, so a review is only as good as its last test. We re-test our top picks every quarter, and whenever a major update lands, and stamp each review with a "last tested" date — so you always see a current verdict, not a launch-day impression.
How We Make Money
Virl earns affiliate commissions when you sign up through our links — at no extra cost to you. That's it. Commissions never change a score or a ranking: our weights are fixed and public, right here on this page. If the best app for you is one we don't earn a cent from, that's still the one we'll point you to.