Our Methodology

How We Test AI Companion Apps

Every app on Virl is rated from 0 to 10 across six weighted pillars — after hands-on testing on accounts we create ourselves — every free tier, and the full signup and checkout flow. No brand buys a score, and no brand edits one.

By Eli Voss, Lead AI Companion Analyst

Where our testing currently stops. Free tiers, character creators, live conversations and checkout pages are all tested first-hand. Paid-tier testing is rolling out app by app — where something sits behind a paywall we have not crossed, we say so rather than score a guess.

Independent

We never accept a brand's account, script, or approval in exchange for a higher score. Rankings are ours, and we'll say when an app we earn from isn't the right pick for you.

Hands-on

Every verdict follows hands-on testing on accounts we create ourselves — every free tier, and the full signup and checkout flow — not a press demo, not a rewritten feature list.

Transparent

The same six-pillar scorecard runs on every app, with fixed public weights. You can check our math on any review.

The Six Pillars We Score

We break every AI companion app into six things that actually decide whether it is worth your money. Each is scored 0–10, then weighted into a single rating.

Pillar Weight What it measures
Conversation & Memory 25% How natural the chat feels, and what it remembers.
Visuals & Media 20% Image realism, character consistency, and speed.
Customization & Personas 15% How deeply you can shape a companion.
NSFW Range & Control 15% What it allows, and how clearly it draws the line.
Value & Pricing 15% What the free tier gives you, and how far money goes.
Trust, Privacy & Billing 10% Discreet billing, data handling, and track record.

How the Overall Score Is Calculated

We don't average the six pillars equally — the things that make or break a companion app count for more. Conversation carries the most weight because a beautiful app with a forgettable chat isn't worth paying for. We multiply each pillar by its weight, add them up, and round to one decimal.

Example

An app scoring 9 on conversation, 9 on visuals, 8 on customization, 8 on NSFW, 8 on value, and 8 on trust earns an overall of 8.5 / 10.

What Our Scores Mean

So an 8 means the same thing on every app we test:

Score What it means
9–10 Best in class. Nothing meaningful to fix.
7.5–8.9 Strong. Minor gaps that don't break the experience.
6–7.4 Solid but limited. Average for the category.
4–5.9 Weak. Real holes you'll work around.
2–3.9 Barely functional.
0–1.9 Broken or missing.

We don't hand out 10s to keep brands happy. A spread of scores is the point — if everything earned a 9, the rating would be worthless.

What Testing Actually Looks Like

Before a single score goes live, here is the kind of thing we put each app through:

Conversation & Memory

We run long, multi-turn chats, drop in a personal detail and check if it is remembered an hour — and a day — later, switch topics abruptly to see if the character holds, and time how long replies take.

Visuals & Media

We generate the same character multiple times to check consistency, push specific poses and outfits to test prompt control, time generation, and count artifacts.

Customization & Personas

We build a companion from scratch to see how many traits we can actually shape, and test voice quality and latency.

NSFW Range & Control

We map exactly where the free-to-paid line sits, and how often the app blocks legal content by mistake. We only ever test lawful adult content — apps with strong safety guardrails score higher on trust, not lower.

Value & Pricing

We track what the free tier really unlocks, what each paid tier adds for the money, and what the checkout page discloses before you commit.

Trust, Privacy & Billing

We check how the charge appears on a statement, what payment methods exist, whether you can actually delete your data, and the app history on leaks.

We Keep Reviews Current

AI companion apps ship updates constantly, so a review is only as good as its last test. We re-test our top picks every quarter, and whenever a major update lands, and stamp each review with a "last tested" date — so you always see a current verdict, not a launch-day impression.

How We Make Money

Virl earns affiliate commissions when you sign up through our links — at no extra cost to you. That's it. Commissions never change a score or a ranking: our weights are fixed and public, right here on this page. If the best app for you is one we don't earn a cent from, that's still the one we'll point you to.

Methodology FAQ

Does Virl get paid to rank apps higher?
No. We earn affiliate commissions when you sign up through our links, but commissions never affect scores or rankings. Our six-pillar weights are fixed and published on this page.
How often do you re-test apps?
We re-test our top-ranked apps every quarter and update the "last tested" date on each review. Major app changes can trigger an earlier re-test.
Do you actually test the adult features?
Yes. NSFW range and control is one of our six pillars. We only test lawful adult content, and we score apps higher — not lower — for having clear, working safety guardrails.
Why score out of 10 instead of 5 stars?
A 0–10 scale gives us the resolution to separate apps that a 5-star system would lump together. The difference between an 8.2 and an 8.7 is real, and it matters when you are choosing.
Who writes the reviews?
Every review is tested and written by Eli Voss, our Lead AI Companion Analyst, based on hands-on use of an account we create ourselves.