A field service software scorecard should make a decision explainable, not manufacture mathematical certainty. Begin with work the company actually performs, weight consequences before meeting sellers, and record the evidence behind every score. This guide provides a reproducible method; it does not publish scores from an executed comparison.

Buyer scenario: polished demos create a tie

Imagine a contractor that has narrowed its search but every presentation looks capable. Dispatch likes one interface, technicians prefer another, finance worries about exports, and the owner is drawn to a broad roadmap. Without a common test, each stakeholder is grading a different promise.

Form a small decision group representing office, field, operations, finance, administration, and security. Define the operating problem in one paragraph. Collect a few anonymized failure patterns: incorrect service address, lost scope change, unassigned return, stale customer update, duplicate inventory use, or failed financial handoff.

Decision criteria and weighting

Create categories for intake, customer and location identity, estimating, scheduling, dispatch, mobile work, evidence, incomplete work, communication, permissions, history, integrations, exports, implementation, training, administration, support, security review, and exit. Write what strong, acceptable, and unacceptable evidence means for each criterion.

Assign weights before demonstrations using frequency, impact, detectability, and recovery effort. Include a small number of blocking requirements that cannot be averaged away. FTC guidance can shape data-security diligence, and PCI DSS can frame payment-scope questions, but actual controls and responsibilities require product-specific evidence.

Reproducible evaluation plan

Give every finalist the same scripted jobs and user roles. Include a routine visit, changed scope, technician reassignment, unavailable part, return work, customer-contact correction, payment or accounting boundary, and export. Ask presenters to distinguish standard behavior, configuration, integration, and roadmap.

Two evaluators should score independently and cite the screen, document, export, or written answer supporting each mark. Reconcile differences after the session. Log unknowns instead of converting them into optimistic scores. This method can be repeated, but no vendor demonstration was run for this article.

Edge case: a favorite fails a blocker

Suppose the preferred interface cannot produce an export that preserves customer-location relationships, or correction of incomplete work erases useful history. Do not compensate with high marks in unrelated categories. Confirm the gap in writing, assess a documented workaround, assign its owner and continuing cost, then apply the predefined blocking rule.

Also test sensitivity: change one reasonable weight and see whether the winner changes. A fragile result signals insufficient evidence or criteria that do not reflect a clear operating priority. Preserve dissent rather than forcing false consensus.

Add an evidence-grade column. Mark whether each answer came from current documentation, a configured demonstration, a sample export, a contract, an implementation statement, or an unsupported assertion. Confidence should affect the decision separately from functional fit. A high score built on unverified promises needs follow-up, not celebration. Record the proposed owner and cost of every workaround, since a workaround without ownership is simply deferred failure. Before finalizing, remove criteria that never influenced a scenario and check that no important exception is represented only by a vague category average.

Run a final calibration session without changing scores immediately. Ask each evaluator to explain one high mark, one low mark, and one unknown using the cited evidence. Resolve different interpretations of the rubric, then rescore only affected criteria. Keep original notes for traceability. A decision group that agrees on a number but not its meaning has not reached a reliable conclusion.

Conclusion: keep evidence beside the score

Choose only after the scripted scenarios, blocker review, implementation model, and reference documents are complete. The score is a summary; the cited evidence and rationale are the decision record. Date the final version and assumptions. A modest scorecard consistently applied is more reliable than an elaborate spreadsheet filled with incomparable impressions.

Traceable evidence

Sources for this decision

2 sources
  1. regulatorData Security guidance for businessesFederal Trade Commission · checked Aug 5, 2026
    Open source ↗
  2. standardsPCI Data Security StandardPCI Security Standards Council · checked Aug 5, 2026
    Open source ↗