How we test
A ranking is only worth as much as the procedure behind it. So we set out what we examine, what we deliberately leave out, and where the limits of our judgements lie.
Five criteria, equally weighted
Every tool is scored from 0 to 5 in five categories. The overall score is the unweighted average. We deliberately avoid weighting, because the right weighting depends on what you do: for a law firm, data protection matters more than value for money; for a start-up it is often the other way round.
Output quality
We give every tool the same tasks from everyday work: summarise a three-page tender, repair a faulty snippet of code, translate a specialist text, generate an image to a precise brief. What we score is how much rework is needed before the result is usable.
Usability
How quickly do people find their way around who bring no AI expertise? We look at comprehensible labelling, a German-language interface, whether settings can be understood, and whether errors are explained rather than merely reported.
Feature set
What counts is what genuinely helps day to day: file handling, context length, export formats, APIs, team features and integration with software you already run. Announcements do not count towards the score.
Value for money
We set list prices against actual benefit and take into account the usage limits of free tiers as well as hidden costs where billing is usage-based.
Data protection
We assess the processing location, the availability of a data processing agreement, the default setting for training on your inputs, retention periods and how comprehensible the vendor's own statements are. This criterion is not a legal opinion but an editorial assessment based on publicly documented information.
What we do not score
We publish no synthetic benchmark results. Those numbers are easy to optimise for and say little about usefulness at work. Nor do we score announcements: features that exist only in a preview count towards the score once they are generally available.
The limits of our judgements
AI tools change fast. A score describes the state of a tool at the time of testing, not a permanent property. Prices are vendor list prices and can change at any time; the vendor's own terms always take precedence. Our data-protection assessment is neither legal advice nor a data protection impact assessment.
Independence and funding
We test with accounts we pay for ourselves. Vendors receive no advance copies of our texts and have no influence on scores or ordering. Should we use affiliate links in future, we will mark them clearly wherever they appear; such a link has no bearing on the score.
Where the results live
Every score is broken down in the corresponding review; the full overview with filtering and sorting is at AI tools compared.
Corrections
Mistakes happen. If you find one, write to us via the contact page. We check every report, correct demonstrable errors promptly and flag substantive changes in the article.
Frequently asked questions about our methodology
Do vendors pay for scores?
No. We take no money for tests, scores or placements, and we do not let vendors review texts in advance.
How is the overall score calculated?
It is the unweighted mean of the five individual criteria, rounded to one decimal place. We deliberately do not weight them, because the right weighting depends on your use case, which is why the individual scores appear on every tool page.
How often do you retest?
Quarterly, and after major model or pricing changes. The date of the test data is in the footer.
What if you get something wrong?
Write to us. We correct demonstrable errors promptly and flag substantive changes in the text.