Privacy scoring

Privacy Scoring Methodology


How SEXTECHGUIDE calculates A-E privacy grades across brands.

Archetype weights

·

Algorithm versions

·

Human review

Each brand we cover is privacy-scored using the framework described here. The framework is conditional: different kinds of brands face different privacy threats, so different scoring categories and weights apply. Every score the current pipeline computes is stamped with the algorithm version that produced it, and changes to this framework do not retroactively alter past assessments: a score that predates versioning keeps its original number, is never recomputed under a newer algorithm, and is shown with its algorithm version marked as not recorded.

Current algorithm version: 2026.12

How a score is built

  1. We classify the brand into a service archetype (see the catalogue below). The archetype determines which scoring categories are applicable AND how each category is weighted, so it changes the number, not just the label. The classification is proposed by an AI model and accepted automatically only when its confidence is at least 0.85; below that it waits for a person to confirm it. A grade is not published at all while the archetype is unconfirmed, because a score produced against a rubric nobody chose is a score with no stated basis.
  2. We scrape the brand's privacy policy and terms of service, plus supplementary policies (cookie policy, GDPR page, security page) when available.
  3. For each applicable category, an AI model rates the brand 1 to 5 against documented criteria and reports its own confidence in the score.
  4. A second AI pass independently fact-checks high-stakes claims (GDPR/CCPA policy-text signals, high-severity concerns) against the source policy text. Where the two passes disagree, the disagreement goes to a cross-provider tiebreaker where that check is enabled, and the field is downgraded to "unverified" only if the disagreement stands. Where the second pass could not run at all, the field still shows as unverified but is NOT demoted: the severity we assigned stands, and our own record says the check could not run rather than that the claim failed. The public marker is the same in both cases, so an unverified label never means, on its own, that a claim was contradicted.
  5. Where the AI reports its own confidence in a category score, we check it against our floor, 0.6 by default. The core categories that apply to the brand's archetype are all required, so if any of them comes back below the floor we produce no score for that brand at all, and the run is held for review rather than published with a gap in it. A category returned with no confidence value is treated as confident and does not trigger this check.
  6. Independent of the AI: we run a Mozilla Observatory security-headers scan and check the brand against the Have I Been Pwned breach database.
  7. The applicable category scores are weighted by the archetype's weight table, which already sums to 100 across the categories that apply, so a category the archetype marks not applicable carries no weight rather than one handed to its neighbours. Where a category that does apply could not be judged, in practice a security scan we could not run, its weight is dropped and the remaining weights are renormalised. That result is modified by transparency and terms-of-service quality, reduced for breaches we have confirmed belong to the brand, and expressed as a score with a letter grade.

How 1-5 ratings become percentages: Each category's 1-5 rating is converted to a percentage on a deliberately non-linear scale: 1 = 30%, 2 = 55%, 3 = 72%, 4 = 87%, 5 = 100%. The scale is chosen so that an "average" rating of 3 lands in the C-grade range, and it is more generous at the low end so a legitimate business with standard privacy practices is not pushed into the bottom grades. The weighted average of these percentages, using the archetype's weight table, is the base score that the modifiers below then adjust.

Service archetypes and their weight tables

Each archetype has a documented privacy threat model. The weight tables sum to 100 across each archetype's applicable categories: N/A categories are excluded entirely (no penalty, no zero contribution).

Archetype Data Collection Data Sharing Data Retention Security Content Privacy
Hardware Product
hardware-product
32% 22% 18% 28% N/A
Connected Hardware
hardware-connected
22% 22% 18% 22% 16%
Subscription Content
subscription-content
20% 26% 22% 16% 16%
Live Interaction
live-cam
18% 22% 18% 14% 28%
Dating / Hookup
dating-platform
22% 30% 18% 14% 16%
Editorial / E-commerce
editorial-ecommerce
26% 26% 18% 22% 8%
Sex-Tech / Health
sextech-health
24% 22% 26% 16% 12%
AI Companion
ai-companion
18% 18% 22% 14% 28%
Support / Advocacy
support-advocacy
14% 24% 20% 18% 24%

Threat-model justifications

Hardware Product
Physical product with no companion app, user account, or cloud sync. Privacy footprint limited to warranty registration and customer-support contact.
Connected Hardware
Physical product paired with a companion app, cloud sync, or biometric/usage telemetry. User accounts present. Common in app-controlled toys.
Subscription Content
Pay-to-watch or pay-to-access content services. Identity, payment, and viewing history are central.
Live Interaction
Cam sites, real-time chat, and live performance platforms. Real-time audio/video and messaging amplify content privacy risk.
Dating / Hookup
Profile-based matching platforms. Photos, location, messaging, and biometric verification are common; data sharing with third parties is a known risk.
Editorial / E-commerce
Review sites, marketplaces, and editorial blogs. Account, payment, and ad targeting are typical; user-generated intimate content is rare.
Sex-Tech / Health
Cycle, fertility, or sexual wellness tracking apps. Health-grade special-category data; potential GDPR / HIPAA-equivalent regulatory exposure.
AI Companion
Conversational AI companions, virtual partners, and intimacy chatbots. The intimate conversation itself is the core asset: chat logs, emotional and sexual disclosures, and their reuse for model training drive the threat model. Persistent accounts and long-lived "memory" are intrinsic. Content privacy is weighted highest, with retention close behind because deletion that does not reach training data is the central failure mode.
Support / Advocacy
Non-commercial organisations: charities, advocacy campaigns, professional and accreditation bodies, universities, and support networks or helplines. No product is sold; the user donates, subscribes to updates, seeks support, or makes a sensitive disclosure. Protecting that disclosure, not over-retaining it, and not onward-sharing it to funders, partners, researchers or authorities is the core threat model.

Scoring categories

Data Collection
What user data is collected, how, and whether collection is minimised. Higher score = less collected, clearer disclosure.
Data Sharing
Whether data is shared with third parties (advertisers, analytics, processors), under what conditions, and with what user controls.
Data Retention
How long data is kept, and whether users can delete it. Indefinite retention without justification scores low.
Security
Independent: based on the Mozilla Observatory scan, which covers security headers and cookie flags. We also run our own direct TLS check; its results are recorded, the HTTPS check is shown on the scorecard, and they feed this score only when the Observatory scan is unavailable. Not derived from the privacy policy.
Content Privacy
How user-generated content (photos, messages, video) is protected. Not applicable to the Hardware Product archetype.

Confidence floor and verification

Per-category confidence floor: We ask the AI for its own confidence in every category it scores, and where it reports one we check it against our floor, 0.6 by default; a category returned with no confidence value at all is treated as confident and is not caught by this check. The core categories that apply to the brand's archetype are all required, so if any of them falls below the floor we do not score the brand at all: no letter is computed, and the run is held for a person to look at, under a hold no editor can simply approve. The withheld run and the categories that caused it are recorded in the assessment audit log. Two things are deliberately different from this: a category the archetype marks not applicable is excluded by design, its weight simply absent from that archetype's table, so nothing is redistributed; and security comes from an independent scan rather than the policy text, so its absence does not stop a score, its weight being dropped and the rest renormalised instead.

Two-pass AI verification: High-stakes claims, GDPR/CCPA policy-text signals and high-severity concerns, are sent through a second AI fact-check call against the source policy text. The second pass returns AGREE or DISAGREE. A genuine DISAGREE that survives the cross-provider tiebreaker, where that check is enabled, demotes the field and marks it "unverified" on the public scorecard. When the second pass simply could not run, whether from a temporary service outage, a missing key, or a source too short to check, the field is not demoted: the severity we assigned stands, and our internal record notes that the check could not run rather than that the claim failed. On the public scorecard both cases carry the same "unverified" marker, so that label never means, on its own, that a claim was contradicted. The claim and the verdict, with the reason for it, are kept on the stored assessment record; where a disagreement goes to the cross-provider tiebreaker, the verdicts are additionally written to the assessment audit log.

Quote grounding: Every quotation and every piece of concern evidence we publish is matched against the source policy or terms. Most appear there word for word. Where the wording differs slightly, we require the quotation's distinctive words to sit together within one short passage of the document rather than being scattered across it. Quotations we cannot match at all, or that turn out to be a geo-block or error interstitial rather than policy text, are discarded before anything is shown.

Publication hold: An assessment is held back from publication and routed to human review, rather than shown automatically, when its own fact-checking flags it: when a majority of the checked claims genuinely failed verification, when the summary was contradicted and a high-severity concern was demoted, or when several quotations could not be grounded. While an assessment is held, the brand's page says a review is in progress and publishes no grade, no score and no category detail. A brand with no published grade should not be read as scoring badly: an assessment is most often held because our own retrieval, analysis or verification did not complete, not because of anything in the brand's policies.

Overall confidence: Each assessment carries an overall confidence rating drawn from the depth of the policy text, the completeness of the AI analysis, the security scan, breach data, and the agreement of the independent second pass. A "high" confidence rating requires a successful security scan in addition to strong policy and analysis coverage; brands whose security surface could not be scanned sit at a lower confidence even when everything else is strong. An assessment whose overall confidence rating falls below our publication threshold, currently 50 out of 100, is held for an editor to review rather than published. That threshold is a separate control from the per-category confidence floor above, which is a different measure on a different scale.

Cross-provider tiebreaker: When the first two AI passes disagree on a high-stakes claim (compliance label, severity), a third pass from a different provider casts a tie-breaking vote. The two-of-three majority verdict decides the public field. All three verdicts, including the disagreement reason from each provider, are kept in the audit log. Cross-provider checks are hard-capped at a small fixed number of calls per assessment so the cost is bounded.

Modifiers and adjustments

  • Policy Quality Modifier (-5 to +5): Adjusts the weighted score based on Terms of Service fairness and overall transparency.
  • Breach Penalty (0 to -20): 3 points per breach we have confirmed is attributable to the brand, plus a severity weighting; capped at 20 points. Matches we have never confirmed contribute nothing.
  • Coverage Bonus (0 to +5): Awarded for publishing supplementary policies (cookie policy, GDPR page, security policy, etc.), capped at 5 points. From algorithm 2026.11, a supplementary document counts toward this bonus only when we actually retrieved its content and that content is substantive. A document whose fetch failed, was blocked (for example by an age gate or bot check), or came back empty or near-empty earns nothing.

Breach history

We check each brand against the Have I Been Pwned breach database. A database match is not by itself evidence that a breach belongs to the brand: names and words collide, and a match on a company name alone can point at an unrelated organisation. Only an attribution we have confirmed lowers a score or appears on the page. Withdrawing a confirmation takes it off the page straight away, but it does not change a score that was produced under an earlier version of the algorithm: we never recompute a stored score against a newer one, so the published letter can keep a penalty from a withdrawn attribution until that brand is next assessed. A breach record naming the same registrable domain we assessed is confirmed automatically; a match on the company name alone is never enough by itself and stays out until a person confirms it. Unconfirmed matches are held back entirely: they are not scored, not counted, and not shown. Where we cannot establish that the brand and the domain we assessed are the same organisation, no breach detail is published for that brand at all.

GDPR and CCPA rows

The scorecard shows what the brand's published policy appears to say about GDPR and CCPA. These rows do not feed the score. They are a reading of the policy text, produced by an AI model and, where possible, checked against a quoted passage from the document itself; they are not a legal determination of whether a company complies with either law, which is not something we are in a position to decide. A row marked as unverified means the label is not backed by a quoted passage that passed our checks. That can happen for more than one reason: we could not match a supporting passage in the document, our independent re-check could not run, or the re-check did not confirm the label. It does not mean the brand is non-compliant.

Letter grades

A ≥ 85 · B ≥ 75 · C ≥ 55 · D ≥ 40 · E < 40

These are the bands used by algorithm version 2026.12, which is the version this page describes. A score keeps the bands it was assessed under: if the bands change, past assessments are not re-lettered.

The bands have not changed since scoring began, so there is no history to show. If they ever do, every past version will be listed here alongside the assessments it applied to.

What a letter can and cannot tell you apart

A score is a weighted average of category ratings that are judgements, not measurements. Two assessors could reasonably differ by a few points on the same evidence, so a difference of a point or two between two brands is not a real difference and should not be read as one. What the letter is meant to separate is a brand near the top of the scale from one near the bottom, not one brand from its neighbour. Because of that, any score landing within 2 points of a band threshold is held for a person to review before it publishes, rather than going live on the letter the calculation produced.

The checks that can stop a grade publishing

A score is computed automatically, and when every check below passes, so is publication: nobody re-reads a clean run before it goes live. An assessment is withheld from the site whenever any of the checks below does not pass. It stays withheld until either a person resolves the reason it was held, or a later scheduled re-assessment of the brand runs and passes every check, in which case the new result replaces the held one and publishes. The software never publishes a held result on its own; only a named person can, through the editorial holds below. The list below is generated from the software rather than written out by hand. It is not the whole of what stops a grade: publication is refused again at the moment of publishing, if no score was computed or if one of these conditions has changed since the assessment ran.

Holds that need the assessment re-run or repaired

Some of these mean the brand's evidence is missing, unusable, or cannot be attributed to it. Others mean our own run failed: it was interrupted, produced no score, or could not persist the source corpus it assessed. No editor can approve past any of them, and each has its own remedy: re-running the assessment, supplying the source text by hand, confirming that the policy we read belongs to the brand, choosing the service archetype, or writing to the company and dealing with any reply.

Snapshot missing · Fetch blocked · Identity unconfirmed · Interrupted run · Incomplete run · Archetype unconfirmed · Notice pending · Manual text · Incomplete data

Holds that need an editorial judgement

These mean the assessment is complete but something about it warrants a person deciding whether to publish. An editor may publish it, and must write down why. That reason is kept permanently against their name.

Low confidence · Fact-check · Grade boundary · Source truncated · Grade changed

Two reviewers for the lowest grade

A grade of E is the strongest adverse assessment we publish about a company, and it requires two different named people to approve it, each giving their own written reason. One person may publish it alone where a second reviewer is not available to them. We do not currently block that route on our own availability check: the software independently checks whether another eligible reviewer existed at that moment, writes what it found, yes, no, or that this could not be established, into the permanent record beside the decision, and lists every single-reviewer publication for later review rather than letting one pass silently.

We tell a company before we publish the lowest grade

Before a grade of E is published, we write to the company, tell them the grade, tell them which document we assessed and when we retrieved it, and give them 10 business days to correct anything we have got wrong. This is a rule that fires on the grade, not a decision we make case by case. If they reply with evidence, publication is held until an editor has dealt with that evidence and recorded what they did about it. We correct on evidence and require nothing in return, and we do not make, withhold or alter an assessment in connection with any commercial arrangement.

Data sources

Anthropic Claude
Reads the privacy policy and terms of service; produces per-category scores 1-5 with confidence values; runs the second-pass verification of compliance and severity claims. The policy documents are fetched afresh for a full assessment, subject to a short content cache of up to 7 days; any change in the retrieved text triggers a fresh analysis. Where the text is unchanged, cached analysis results may be reused for up to 30 days, and verification verdicts for up to 7 days.
OpenAI
Cross-provider tiebreaker. Only invoked when the primary 1st and 2nd verification passes disagree on a high-stakes claim. Returns AGREE / DISAGREE; the two-of-three majority decides. Hard-capped at a small fixed number of calls per assessment. The specific OpenAI model used for each assessment is recorded in the per-assessment audit log; we refresh model selections periodically as model capabilities improve.
Perplexity
Discovers the canonical URLs of the brand's privacy policy, terms of service, and supplementary policy pages. Used only for URL discovery, not scoring.
Mozilla Observatory
Scans the brand's domain for security headers, cookie flags, and related web-security tests. Free public service. Drives the Security category score; TLS and certificate checks are made by our own direct scan of the domain.
Have I Been Pwned (HIBP)
Checks the brand against the public breach database. Drives the breach penalty modifier. Cache refreshed weekly.

Policy discovery and scraping

To find a brand's published policies we use, in order: any URLs an editor has supplied; links found on the brand's own homepage and in the sitemap files its robots.txt advertises; and a set of common policy paths. If a privacy policy or terms of service is still missing after that, we ask an external search service for candidate addresses and check anything it returns the same way we check candidates from any other source. Requests to a brand's site carry a user agent that names this project and links back to it, and our policy fetching is paced: successive checks against the same host normally wait at least a second apart.

We fetch published policy pages only. The fetch surface is the brand's own domain and its subdomains, plus a short allowlist of well-known consent-management and policy-hosting providers that some brands use to publish their notices (for example OneTrust, Cookiebot, Termly, iubenda, TrustArc, Osano, Usercentrics and Didomi), and the Internet Archive for the fallback described below. We read robots.txt only to discover the sitemap files it advertises; we do not evaluate its crawl-exclusion (Disallow) rules. Every request to any of these hosts carries the same identifying user agent. Finding a policy can mean checking many candidate addresses; the number of checks in an assessment is capped, the checking is spaced out in the same way, and only pages that prove to be policy documents are kept.

Some brands place their policies behind an age-verification wall. Because these are adult-industry services whose published legal documents we are reviewing for accountability, we take a limited, transparent set of steps to reach the public policy behind such a wall: we may present a standard age-confirmation cookie (asserting the reviewer is over 18), and, if a wall still blocks the page, we may follow a single "enter" or "I am 18+" link on the wall. This traversal is GET-only; submitting forms by POST is disabled by default, and we never sign in or submit account credentials. We only ever score a re-fetch of the brand's own policy page reached this way, never the wall or the landing page itself.

When automated retrieval cannot reach a policy at all, we fall back in two ways. First, we look for the most recent working copy of the page in the Internet Archive (Wayback Machine) and, where a suitable snapshot exists, assess that archived copy. Second, an editor may paste the policy text obtained by hand from the brand's own site. In both cases the source and its date are recorded, and where an assessment was produced from an archived copy the scorecard says so.

Framework and policy changelog

Changes to how we decide what to publish, as opposed to how a score is computed. These do not change any stored score. The algorithm changelog below covers changes that do.

22 August 2026
Published this section, the breach-history and GDPR/CCPA notes above, the human-review section, and the statement of what a grade is. The grade bands and the 1-to-5 scale on this page are now rendered from the code that applies them rather than written out separately, so the page cannot describe thresholds the software does not use.
22 August 2026
A grade of E now requires two different named reviewers, and we write to the company before publishing one, giving them ten business days to correct us. A reply containing evidence stops publication until an editor has dealt with it. Where a second reviewer is genuinely unavailable, one person may publish alone, but that is recorded as a single-reviewer decision.
21 August 2026
Publishing a held grade now requires a written reason, recorded permanently against the name of the person who published it. An assessment that fails several publication checks now records all of them rather than only the first, and every one of them has to be resolved before it can publish.
21 August 2026
Assessments now record which rubric produced them, and a grade is withheld where the brand has no confirmed service archetype. Where we cannot establish that the policy we read belongs to the brand, the whole grade is withheld rather than just the breach detail.

Algorithm version changelog

Changes to how a score is computed. An assessment keeps the version it was produced under and is not recomputed when a newer version ships.

2026.12 (current)
Content privacy rubric rewritten. The rubric now asks the AI to judge each checkpoint from what the published documents state or omit, and says plainly that where a subject applies and the documents do not address it, that silence is itself a finding rather than a gap in our evidence. The confidence instruction was rewritten the same way and applies to every category, not only content privacy, so a score resting on documented silence is now reported as well supported rather than as uncertain. One checkpoint was dropped from each version of the rubric, taking both from five to four. Grade bands are unchanged. Existing assessments keep the version they were produced under and are not recomputed.
2026.11
Coverage bonus now counts only supplementary policy documents whose content was actually retrieved and substantive; failed, blocked, or empty fetches earn nothing. Every assessment also now persists an immutable evidence snapshot (source URLs, retrieval method, fetch times) before it can publish.
2026.10
Cross-provider verification reliability fix. When the third, cross-provider tie-break check that resolves a disagreement between the first two AI passes cannot run, the claim now defers to the two-pass verdict (treated as unverified) instead of defaulting to verified, so a transient verification failure can no longer silently keep a claim the second-pass fact-check rejected. Existing archetypes, weights and stored scores are unchanged.
2026.09
Added the Support / Advocacy archetype for non-commercial organisations: charities, advocacy campaigns, professional and accreditation bodies, universities, and support networks or helplines, where a user's sensitive disclosure or support-seeking is the core asset rather than an account and ad targeting. Its weight table leads with content privacy (24) and data sharing (24), reflecting that protecting the disclosure and not onward-sharing it to funders, partners, researchers or authorities is the central threat. AI rubric refinements cover support-seeker anonymity and disclosure confidentiality, retention and deletion of support/disclosure records, and onward sharing with funders, partners, researchers and law enforcement. Existing archetypes and their stored scores are unchanged.
2026.08
Breach attribution hardened. A known data breach now affects a brand's score, and is shown publicly, only when it is matched to the brand by its registrable domain or has been editorially confirmed; name-only matches are treated as unconfirmed and excluded from the breach penalty (and from the public breach history) until reviewed. The breach severity penalty is computed from the confirmed breaches only. This prevents a brand whose name merely appears within an unrelated company's breach record from being penalised or labelled. Other archetypes, weights and category scoring are unchanged.
2026.07
Added the AI Companion archetype for conversational AI / virtual-partner / intimacy-chatbot brands, with a weight table led by content privacy and retention, and AI rubric refinements covering conversation-log handling and training-data reuse. Existing archetypes and their scores are unchanged.
2026.06
Per-archetype weight tables. Evidence-grounded policy-text signals replace legal-sounding compliance labels. Two-pass AI verification of policy signals and severity claims treats verifier failures as unverified. Source snapshots preserve the text used for each assessment.
2026.04
Per-archetype weight tables. Two-pass AI verification of compliance and severity claims. Confidence floor (0.6) per AI category. Per-archetype rubric refinements for dating platforms and sex-tech / health.
2026.03
Service-archetype framework introduced; N/A categories excluded from scoring; AI archetype classifier added; methodology page launched.
pre-2026.03
Shared fixed weight set across all brand types. No archetype awareness, no two-pass verification, no confidence floor. Assessments produced under this version retain their original score and are not retroactively recomputed under newer algorithm versions.

Limitations and disclaimers

What a grade is

A grade is our editorial opinion of how a company treats personal data, formed by reading the documents that company publishes about itself and scoring them against the rubric on this page. It is an assessment of published policy and observable security configuration. It is not an audit, not a certification, and not a finding of fact about anything the company does internally, which we cannot see. A high grade means the published policies read well against our criteria; it is not a guarantee of good behaviour. A low grade means those policies read poorly against our criteria; it is not an allegation of wrongdoing.

For information only. Privacy scores are published as editorial information to help readers compare brands. They are not personal advice, a recommendation to use or avoid any service, a warranty of any brand's privacy practices, or a guarantee against future incidents. A high score does not mean a brand is safe; a low score does not mean a brand is unsafe. Decisions about which services to use are the reader's own, and readers should always review a brand's current privacy policy and terms of service before signing up.

Not legal or compliance advice. Privacy scores are editorial assessments of publicly available information. They are not legal advice, regulatory determinations, certifications, or audit results. GDPR and CCPA fields are policy-text signals only: they describe whether the published policy appears to address specific topics, not the brand's actual regulatory standing.

Public information only. Scores reflect what a brand publishes. Private practices, internal contracts, undisclosed data-handling, and operational behaviour not documented in the privacy policy or other public materials are not assessable and are not included in the score. Opacity itself is captured in the transparency component of the score.

Scores age. Each assessment is timestamped on the scorecard. Brands change their policies, security posture, and infrastructure; an assessment reflects what was true at the time it was generated. Refreshes are scheduled on a rolling basis but are not instantaneous, and a score may be out of date by the time you read it.

Commercial relationships

SEXTECHGUIDE may earn affiliate commissions from some of the brands it assesses. Privacy scores are produced by the automated methodology described on this page and are not influenced by whether a brand is a commercial partner: a commercial relationship does not raise a brand's score, and the absence of one does not lower it. Where a brand is linked commercially, that relationship is disclosed at the point the link appears.

Disputes and corrections

If you are a brand and believe an assessment is inaccurate or out of date, contact us via the "Request a correction" link in the scorecard footer, or via our published contact page. We respond to correction requests within 21 business days and will either: republish the scorecard with the change applied, mark the assessment as under review pending re-scoring, or explain in writing why the score stands. We log every assessment with timestamps and source attribution and can review specific category scores or the archetype classification on request.