Guide
AI App Audit Report: What It Should Include and How to Use It
Learn what an evidence-based AI app audit report should contain, how to read scores, and how to prioritize findings.
An AI app audit report is a report about an AI-built public website or web app, not an evaluation of an AI model. It should explain what a visitor can understand and do at an authorized public URL, support findings with traceable evidence, and help an owner choose the next fix. It is not automatically a code review, prompt-history review, security assessment, or inspection of private records.
State the scope before the score
Open with the exact public URL, review date, environment, available build or version, viewports, and the pages and actions included. Record exclusions such as login, live payment, irreversible deletion, or a form that could not be submitted safely. Scope makes every later claim legible. Without it, a reader cannot tell whether “works” means one observed public path or the entire product.
Is My App Slop's automated $27 product reviews a public website or web app. It can follow public links and buttons and interact with eligible public forms. It does not enter behind-login areas, inspect source code, read prompt history, or verify private database records. It is not exhaustive quality assurance.
Use the six public product dimensions
A consistent report reviews six dimensions. Clarity and promise asks whether a new visitor can identify the product, audience, outcome, and next step. Craft and design covers hierarchy, legibility, consistency, complete states, and details that affect comprehension. Trust covers visible identity, pricing, terms, support, privacy information, proof, and limitations. Flow and usability follows the main public path and its recoverable errors. Mobile and access checks phone-width behavior, zoom, labels, focus, reading order, and keyboard basics without claiming accessibility compliance. Conversion follows the path from promise to the requested or paid action without completing an unauthorized live charge.
These categories keep a report centered on the product a visitor actually encounters. A separate specialist can inspect code maintainability, private permissions, model behavior, or legal obligations when that work is authorized.
Treat scores as prioritization opinions
An overall score and category scores summarize the auditor's judgment under a stated rubric. They are not objective certification, a warranty, or proof that every path works. Explain what pushed a category up or down. A single broken primary path may deserve more weight than several polished secondary pages.
Compare scores over time only when the URL, scope, rubric, and relevant actions remain comparable. If a later review adds login, payments, or a different device set, describe it as a broader review rather than implying a precise trend from unlike evidence.
Separate coverage from outcome
“Tested” and “not tested” describe coverage. “Pass” and “fail” describe the result of a performed check. A payment flow excluded to avoid a live charge is Not tested, not Pass. A visible button that led to the expected public page is Tested/Pass. A form that showed a confirmation but lost its result after reload is Tested/Fail. Keeping both pairs prevents unknown areas from inflating confidence.
Separate observations from inferred causes
The report should say “the confirmation appeared, then the submitted item was absent after reload” when that is what happened. It may add “the cause could involve persistence” as an inference, but should not declare an unseen database fault. Public evidence supports observations about public behavior. Root-cause claims require access and investigation that may not be part of the product.
Make each finding traceable
Use a table that connects the visitor's action to a repair and later retest:
| Page/action | Expected | Actual | Evidence | Impact | Proposed fix | Retest |
|---|---|---|---|---|---|---|
| Hypothetical booking form; submit valid controlled details | Clear confirmation and a retained reference | Confirmation appears, but the reference is absent after refresh | Before/after screenshots and the recorded public steps | Visitor may believe a request exists when it cannot be retrieved | Preserve the existing form and confirmation copy; make the successful result persist and expose a reference | Not retested |
| Hypothetical pricing CTA; activate “Start booking” | Booking step opens | CTA returns to the same section | Screenshot and destination URL | Primary conversion path stops | Point the existing CTA to the booking step and retain analytics | Pass on corrected build |
Both rows are explicitly hypothetical examples, not Is My App Slop findings or customer outcomes. A real report should attach only the minimum evidence needed to reproduce the behavior and should avoid personal data.
Rank repairs by consequence
Lead with data exposure, wrong charges, lost work, misleading claims, and a broken primary journey. Follow with repeated recoverable friction, then visual polish. Severity should express likely consequence, not how dramatic the screenshot looks.
A narrow repair prompt for the first hypothetical row could be: “Fix persistence for a successfully submitted public booking. Preserve the current fields, validation copy, confirmation design, analytics, and unrelated routes. After the change, repeat the same controlled submission, reload, and record whether the booking reference remains.” The prompt identifies one failure, preserves constraints, and states the retest. It does not ask for a broad rebuild that could change unrelated behavior.
Show what was not tested
Limitations should name unavailable logins, time limits, excluded side effects, inaccessible dependencies, and evidence that could not be collected. An automated public audit does not inspect code, private access rules, source prompts, internal logs, or database records. It also cannot establish legal compliance, security, complete accessibility, or usability with representative real people.
Those can be valuable separately scoped reviews. A security specialist may inspect threats and technical controls. An accessibility specialist and disabled participants can provide deeper evaluation. A developer may review source quality and architecture. A product researcher may observe target users. Keeping those options separate makes the public report more credible, not less useful.
Use the report as a working baseline
Turn each material finding into an owner, repair, evidence requirement, and retest status. After changes, repeat the same page, action, expected result, and rubric. Do not change “Not retested” to “Pass” because code changed; change it only after the behavior is observed again.
Share the prioritized summary with decision-makers and the detailed evidence with the people making fixes. The real sample report shows how this product presents its public findings without inventing a score here. The Lovable app audit guide explains the six-dimension review, while the testing workflow covers deeper owner-controlled verification.
A useful report does not claim certainty it did not earn. It states the public surface reviewed, distinguishes observation from inference, marks unknowns honestly, and gives the owner a reproducible next action.
