BlogCRO

B2B Website Performance Scorecard: 12 Evidence Checks

A practical worksheet for deciding what to improve, what to measure and what to leave alone.

Abstract blue dithered field with the centred headline What should you fix?

A B2B website performance scorecard helps a marketing team connect website observations to decisions. It should show who the site attracts, whether buyers can evaluate the offer, where the enquiry process needs work and which changes have enough evidence to justify investment.

Use the 12 checks below to review a homepage, a priority service or product page, and the path to an enquiry. Each check asks for an evidence record and a next action. The result is a list of priorities you can discuss with sales, your website team and leadership.

The scorecard works for a site on Webflow or another platform. It also works when traffic is too low for a useful A/B test. Conversion rate optimisation, or CRO, includes research, measurement and usability repairs as well as controlled experiments.

How do you use the B2B website performance scorecard?

Choose one audience and one commercial outcome before reviewing pages. For example: heads of operations at software companies evaluating an implementation partner, with a qualified discovery call as the outcome.

Record the period, the pages and the evidence available. Start with complete weeks or months so that a partial day does not distort the comparison. Note campaigns, outages or tracking changes that could affect the figures.

For each check, use one of four statuses:

  • Supported: current evidence meets the requirement you wrote down.
  • Gap: a specific observation shows an unmet requirement.
  • Unknown: evidence is missing, too old or too thin.
  • Not applicable: the check does not fit this purchase or page. Record why.

Count the statuses to show how much of the review is supported. Keep them separate. We do not turn the count into a revenue prediction or an industry percentile. A broken form deserves action even when every other check looks healthy.

Use this record for each finding:

Page and audience:
Requirement:
Status:
Evidence, source and date:
What remains uncertain:
Proposed action:
Owner and review date:
Measure of success:

The worksheet is the working document. Keep links to screenshots, analytics reports, research notes and CRM records beside the relevant entry.

Can buyers understand and assess the offer?

These first four checks examine the decisions someone needs to make before contacting you. An early Nielsen Norman Group study of B2B usability found that research and multi-person purchasing created information needs extending beyond a transaction. That research dates to 2006. Use it as a reason to investigate your own buyers’ tasks, with current customer evidence guiding the details.

1. The page identifies a buyer and a problem

Ask a person who fits the audience to read the page and explain who it serves, what it helps them do and when they would consider it. Write down their interpretation before explaining your own.

Evidence to collect: the exact words on the page, the participant’s explanation and the part that caused uncertainty.

A gap looks like: a buyer can name the broad category but cannot tell whether the service supports their company size, use case or implementation needs.

Next action: clarify the audience or use case in the opening copy, then repeat the task with another relevant buyer. A colleague’s preference alone does not establish comprehension.

2. The page answers a real evaluation question

Review recent discovery-call notes, sales objections and lost-opportunity reasons. Select one question that buyers need answered at this point, such as what a migration includes or how an existing integration will work.

Evidence to collect: where the question came from and the passage that answers it.

A gap looks like: the answer exists in sales conversations but is missing from the product or service page. Another common gap is an answer that requires readers to understand internal terminology.

Next action: publish a direct answer in the relevant context. Give technical detail its own supporting page when the decision requires it. Link to that detail where the question arises.

3. Claims have relevant proof

Match each important claim to evidence of the same outcome. Delivery examples support delivery capability. A page-speed result supports a performance claim. An experiment about button clicks supports a claim about those clicks.

Evidence to collect: the source, metric, period and conditions behind each claim. Record public-use permission for customer material.

A gap looks like: a headline promises revenue growth while the accompanying case reports traffic or engagement. A logo alone also leaves the reader to infer what work was done.

Next action: narrow the wording to what the evidence demonstrates, or gather the missing evidence. The B2B website trust guide covers how to present proof alongside the offer.

4. The next step has a clear commitment

Read the CTA and destination together. Can a buyer tell what they will receive, what information they need to provide and what happens after they act?

Evidence to collect: the button label, destination heading, required fields and confirmation message.

A gap looks like: “See the platform” opens a sales qualification form with no explanation. A resource invitation that unexpectedly asks for project budget is another mismatch.

Next action: describe the actual next step beside the CTA. Give someone still researching a useful route, such as a public checklist or worked example. Reserve detailed qualification for a request that needs it.

Where should CRO measurement start?

Start with the difference between a visit, a submitted form and a qualified opportunity. Checks five through eight establish whether the numbers describe progress through the intended process.

5. Traffic sources can be compared on a consistent basis

Separate major sources before interpreting a change in the site-wide conversion rate. Record the landing page and campaign where available.

Evidence to collect: visits and enquiries by source, together with the attribution method and the share with missing source data.

A gap looks like: a large referral spike lowers the overall enquiry rate, and the team concludes that the website has become less effective. The audience mix may have changed.

Next action: compare the established sources separately, then assess the new source on its own evidence. Google Analytics distinguishes user, session and event acquisition scopes; choose the scope deliberately. Use the same definition in each comparison.

6. The enquiry path has measurable stages

A useful lead-generation path might contain page view, form start, valid submission and sales acceptance. Use the stages your system can reliably capture.

Evidence to collect: distinct visits or people reaching each stage, the counting rule and a successful test record.

A gap looks like: the dashboard reports hundreds of CTA impressions alongside a small number of enquiries, without distinguishing exposure from action. Repeated form events can also inflate totals.

Next action: document the event definitions and verify one complete journey. Test retries and errors too. Reconcile successful submissions with the actual inbox or CRM before using them in a business report.

7. Sales qualification reaches the website review

Agree what counts as a relevant enquiry. Include the reason for excluding supplier pitches, recruitment requests, duplicates and enquiries outside the service scope.

Evidence to collect: received date, source where known, qualification status, reason and review owner. Keep personal details in the appropriate private system.

A gap looks like: a campaign creates more submissions but the sales team accepts fewer useful conversations.

Next action: add a short qualification review to the reporting routine. Preserve both submitted and accepted counts. This lets the team see whether a change improved volume, quality or neither.

8. Form errors can be diagnosed

Try the form on a phone and with keyboard navigation. Trigger a required-field error, correct it and complete the journey. Check that the system retains valid answers and confirms success clearly.

Evidence to collect: the device, the step, the error, whether recovery worked and whether the record arrived.

A gap looks like: the form says an answer is invalid without identifying the field or explaining how to fix it.

Next action: repair the observed fault and verify it again. The W3C guidance on form notifications describes ways to communicate errors and successful completion accessibly. Four errors in analytics, without a reproduction or explanation, would be a finding to investigate.

Are visibility and delivery being measured usefully?

The final checks connect acquisition, site quality and implementation to decisions. They also expose where more research is needed.

9. Search and AI visibility have named questions and dates

List the commercial questions the site needs to answer. Link each question to an existing page, any search evidence and dated observations from AI answer checks.

Evidence to collect: query or prompt, market, engine, date, relevant page, observation and whether a valid answer was returned.

A gap looks like: a combined AI visibility percentage appears without the questions or denominator behind it.

Next action: record the observation set and preserve missing checks. Keep citations, referral visits and enquiries as separate measures. Google’s guidance for AI search features also places eligibility within its wider Search requirements; a checklist cannot guarantee inclusion in an answer.

10. Performance evidence reflects actual use

Identify whether reported performance comes from a laboratory test or real users. Check the priority page and device, then look for a specific usability consequence.

Evidence to collect: tool, date, page, device, available field sample and the observed issue.

A gap looks like: a single desktop test is treated as proof that the mobile form works well. Another gap is a low performance score with no investigation of which part of the experience is affected.

Next action: reproduce the slow or unstable interaction and investigate its cause. Web Vitals documentation distinguishes field measurement from diagnostic tools. Its thresholds help assess performance; they do not estimate the additional leads a repair will generate.

11. The proposed test has a feasible sample

Define the outcome, the eligible audience and the smallest change worth detecting. Estimate the sample and duration using the method supported by the experiment tool.

Evidence to collect: baseline rate, eligible traffic, chosen statistical settings, allocation, required sample and business deadline.

A gap looks like: the test needs several months of eligible visitors, but the team expects a decision next week.

Next action: change the research method or the decision scope. Buyer interviews, task-based usability sessions and a verified defect repair can still inform work. Optimizely’s explanation of minimum detectable effect is a useful starting point for sample planning.

12. Every priority has an owner and acceptance check

Turn findings into work that someone can finish and verify.

Evidence to collect: a named owner, the precise change, dependencies, delivery date and acceptance check.

A gap looks like: “Improve conversions” remains on a task board for a month because nobody can tell what completion means.

Next action: write a concrete action such as “Explain the discovery-call format beside the form; verify the copy on phone and desktop; review qualified enquiries after two complete weeks.” Note that a before-and-after comparison cannot isolate the effect of the copy from other changes.

How do you choose the first improvement?

Handle a reproduced failure that blocks a necessary task first. Then consider how many relevant buyers encounter a gap, how directly it affects their decision and how strong the evidence is. Estimate effort and dependencies before committing.

Keep an “unknown” entry in the research queue when it could change a substantial investment decision. An unknown source split might justify a measurement task before a redesign.

Use this sequence in the review meeting:

  1. Confirm the commercial outcome and the audience.
  2. Check whether any critical journey is broken.
  3. Choose a buyer question or task with a documented gap.
  4. Agree the smallest useful change or research activity.
  5. Record who owns it and how the result will be assessed.

Avoid adding the same symptom to several workstreams. A confusing offer can affect hero copy, paid landing pages and forms; record the shared question once and specify where the answer needs to appear.

A worked example for a B2B service website

Consider this illustrative month: 2,000 sessions, 80 sessions viewing the contact page, 12 submitted enquiries and four accepted opportunities. Sales has reviewed all 12 enquiries. A directory referral supplied 1,200 sessions and one submission, which was a supplier pitch. Other sources supplied 800 sessions, 11 submissions and all four accepted opportunities.

The overall submission rate is 0.6%. The directory’s rate is approximately 0.08%; the other sources’ rate is approximately 1.38%. These are arithmetic descriptions of this invented example, with sessions as the denominator. They are not industry benchmarks.

The review would record a source-quality finding and investigate how buyers reach the enquiry page. It would also ask what prevented the eight unaccepted submissions from qualifying. Changing button colour has no supporting evidence here.

Suppose a usability session then shows that a relevant buyer cannot tell whether ongoing support includes development. That produces a specific copy task. The scorecard entry can carry the quote, affected page, owner and acceptance check without pretending that the observation predicts an uplift.

Frequently asked questions

What is a good B2B website score?

The useful result is a supported set of decisions. This scorecard counts supported checks, gaps, unknowns and exclusions separately. It does not assign a universal pass mark. The importance of each gap depends on the task it blocks and the buyers it affects.

Can we use this before a redesign?

Yes. Use it to identify which constraints a redesign must solve and which parts already work. Attach the findings to the brief, with baseline measures and an owner for each requirement.

Does every CRO recommendation need an A/B test?

A/B testing is useful when a controlled comparison can answer the decision at the available sample size. Reproduced defects, accessibility failures and missing measurement can be addressed and verified directly. Record the type of evidence used and avoid claiming experimental lift from an untested change.

Can Rialto supply the evidence?

Rialto can supply saved behaviour, acquisition, content and AI-answer observations where the project has coverage. Check their dates and definitions. Qualification, buyer interviews and delivery decisions still require evidence from the team and its systems.

For help turning the completed worksheet into a measured work plan, discuss your website priorities. Our CRO service and measurement support explain how that work is run.