Consistent criteria
Every app is measured against one published model, so a 7.0 means the same thing on every listing.
Our review process
Every NSFWRating score uses the same weighted model across 10 criteria, chat quality, privacy, memory, media, customization, free tier, pricing, safety, mobile and user sentiment, so 121 apps can be compared on identical terms.
What guides every review
The same rules apply to every product, regardless of popularity, commission rate or commercial relationship.
Every app is measured against one published model, so a 7.0 means the same thing on every listing.
Claims must trace back to hands-on testing, the vendor's own policy pages, or reputable public sources.
Scores reflect our own evaluation, not unverified star averages, and not paid placement.
Pricing, limits, features and privacy terms are re-verified, and each listing records when it was last checked.
The weighted scoring model
The overall score is a weighted average of 10 criteria. Higher-impact areas such as chat quality and privacy carry more weight than smaller usability factors. The weights below are the live values used by the scoring engine and total 100%.
What we test
Each category has a defined set of checks. Reviewers record evidence, note limitations, and assign a criterion score before the overall result is calculated.
We assess coherence, repetition, tone control, refusal behaviour, realism and response consistency across repeated sessions.
We review account and chat deletion, what signup demands, data-training rules, and how clearly the privacy policy states them.
We test how well an app holds on to preferences, character detail and conversation context between messages and sessions.
We look at image and video output, generation controls, wait times, consistency, and the limits attached to each plan.
We review character builders, persona controls, scenarios, saved presets and how deep the available adjustments go.
We verify daily limits, starter credits, which features are locked, trial length, signup friction and any hidden restrictions.
We check whether plans and credit costs are visible before signup, plus refund rules, subscription terms and renewal behaviour.
We examine age gates, consent rules, fictional-only boundaries, moderation, and the policy on uploading images of real people.
We test responsive behaviour, mobile-web usability, app availability, navigation speed and interaction stability.
We use moderated, aggregate public sentiment as a supporting signal only, filtering out spam and unsupported claims.
How we gather data
Hands-on evaluation is combined with vendor documentation and reputable public sources. Every material claim should connect to something observable.
Confirm product scope, plans, policies, supported platforms and current claims.
Use the product across repeated sessions, recording results against each criterion.
Pricing, privacy, feature and safety claims are checked against the vendor's own pages.
Criterion scores are weighted and converted into the final score out of ten.
Listings are revisited when plans, limits, features, policies or quality change.
Editorial independence
Featured placements and affiliate relationships never change a product's editorial score. Disclosures stay visible, and criticism is not removed for commercial reasons.
The evaluation is completed against the criteria above, independently of advertising, commissions or partnership terms.
Read the affiliate disclosure →Featured positions are identified wherever they appear and are never blended into a score calculation.
How featured placement works →Limitations, policy concerns and pricing weaknesses stay in the review even when a brand advertises with us.
Report a correction →How often we update
Prices, free-plan limits, features and policies move constantly in this category. High-impact listings get priority, and every listing shows the date it was last checked so you can judge how fresh it is.
Common questions
How the scores are produced, what they include, and what they deliberately do not.
No. Scores are editorial, calculated from the weighted criteria above. Moderated public sentiment contributes only 5% of the total.
No. Scores are determined independently of commissions, sponsorships, coupon agreements and advertising. Paid placements are labelled and kept separate from the editorial score.
Hands-on testing is the standard for a full review. Some listings are scored from the vendor's own documentation and observable public evidence instead; where that is the case the review says so rather than implying a deeper test.
Rankings move when pricing, free-plan limits, features, privacy terms or safety rules change - or when a competing product improves. Each listing records when it was last checked.
It is a weighted average out of 10 across 10 criteria. Across the 121 apps scored today the range runs from 3.3 to 7.6, with an average of 6.2 - so a 7 is genuinely strong rather than average.
Send a factual correction with supporting evidence through the report form. Corrections are reviewed editorially and never require a commercial relationship.
Browse every reviewed product, check category rankings, and inspect privacy and pricing before you choose.