Research methodology

How We Test and Evaluate Field Service Software

How we test workflows, verify changing product facts, label evidence, handle incomplete research, and keep commercial terms out of recommendations.

On this page

This page explains what we test, how we label evidence, when a product is eligible to be ranked, and what happens when the evidence is incomplete. It describes the public rules readers need to interpret our reviews; product-specific run details live with each review’s test record.

Method at a glance

  1. We evaluate real service-business workflows, in the product, rather than feature lists.
  2. We use comparable tasks and starting conditions where a meaningful comparison is possible.
  3. We separate first-use friction from repeat-use performance.
  4. We label hands-on, vendor-documented, reported and editorial claims apart, and never blend them.
  5. We verify changing facts and show the date each was checked.
  6. We publish a ranked order only after the evidence and human-review gates are met.
  7. We exclude commissions and commercial terms from eligibility, evaluation, order and verdict.

What we test

A feature list tells you a product has “scheduling”; it does not tell you how many screens it takes to reschedule a job when a customer calls from the driveway, whether the estimate you sent last week turns into that job without re-entering it, or whether the plan you can afford includes the part you need. So we test workflows.

The core service workflow. The jobs common to most service businesses: taking in a new customer, estimating or quoting, scheduling, running and documenting the job, invoicing, taking payment, customer communication, and the follow-up that turns one job into repeat work. Each product gets comparable tasks and starting data where that comparison is meaningful; product limitations, account access, market differences and vendor rules can prevent literal identity, and where they do, the page says so.

The trade-specific workflow. Only the criteria that materially change the purchase decision for a trade. A chimney business may need property and chimney-system history, inspection documentation and repeat-service reminders; a painting contractor may put estimating depth, multi-day crew work and job costing first; a detailer may need vehicle records and mobile booking with deposits. A product that handles the shared workflow well and the trade’s own work badly is not a fit, and the pages for that trade say so.

Not every product or market enables every test action. A test we could not run is labelled incomplete; it is never treated as a failure and never scored.

When a test step could trigger a real-world consequence, such as taking a payment or sending a reminder, we use controlled synthetic data and require explicit human approval.

Evidence labels and testing statuses

Every material claim on this site carries one of four evidence labels. They appear on reviews, pricing guides and trade guides exactly as they appear here.

Label What it means What it does not mean
Hands-on verified Observed in a preserved test run by a person, or confirmed directly in the product interface on a stated date. Not a promise that every workflow was tested, and not a claim about how the product feels in daily use.
Vendor-documented Supported by current first-party vendor material: documentation, pricing pages, terms or help content, with the source and the date it was checked. Not independently proven performance: it is what the vendor publishes, checked on a date.
Reported Attributed to a named third party: a user, a reviewer, a forum or another external source. Never presented as our own finding, and never treated as universal.
Editorial assessment Our human-reviewed interpretation, reasoned from the cited evidence and labelled as ours. Not an objective fact, and never stronger than the evidence it is reasoned from.
The same four labels appear on reviews, pricing guides and trade guides exactly as they appear here.

A label applies to the claim it sits beside. Where every claim in a bounded section, a list or a table column rests on the same kind of evidence, the label is stated once for that section or column; where the claims differ, each carries its own. A recommendation is always labelled as our assessment, even when the feature fact behind it is vendor-documented.

A product page also carries a testing status: how far the product is through our standardized suite. A status describes progress, not the strength of any one claim, and “vendor-documented” is never a testing status.

Not yet tested
No workflow scenario has been completed hands-on. Product facts on the page are labelled by their own evidence.
Testing in progress
Some core workflow scenarios are complete hands-on; the page prints how many, out of how many, and the latest test date.
Core testing closed
Hands-on testing ended before every core workflow scenario was completed. The page prints how many are complete, the date testing closed and why, and names each scenario not completed; the completed findings stand with their dates, and the gaps are named, not filled.
Core workflow complete
Every core workflow scenario has been completed hands-on in a preserved test run.
Trade-specific testing complete
The core suite and the trade-specific scenarios for the page’s trade are complete.
Retest required
Hands-on evidence is past its window or the product changed materially; findings stay visible with their dates until retested.

Two further states can appear beside a single fact: Not yet tested when a workflow has not been completed under our testing and no conclusion about it is published, and Recheck required when a time-sensitive fact is past its re-verification window and is being rechecked.

How testing and verification work

A person runs the hands-on workflows, in an ordinary browser, on an authorized account, with controlled synthetic data. Where the difference matters, a first-use pass that measures how findable a path is and a repeat pass that measures the workflow once it is known are recorded separately. Hands-on sessions are documented, with recordings or other preserved evidence where the method requires it, so findings can be reviewed and material claims traced back to what supports them.

Changing facts use the most appropriate current source. For workflow behaviour, direct observation in a preserved test of our own is usually strongest. For current pricing, plan limits and formal product policies, current first-party documentation may be controlling. Vendor support can clarify ambiguous documentation, and third-party reports are used for attributable real-world experience we cannot establish from first-party sources alone. When vendor sources disagree, we identify the controlling source where its scope and authority resolve the question and explain our choice. If the disagreement remains unresolved, we do not present the disputed detail as settled. A claim nothing supports is withheld, not estimated.

A human executes product interactions and approves subjective conclusions. Software, including AI, may assist with transcription, organizing records, consistency checks, calculations, drafting and quality checks, but it is never presented as a person’s experience of the product and it does not approve a recommendation. Words such as “intuitive” or “frustrating” describe an experience; they appear only when the person who ran the session recorded that judgment as part of the evidence, and we do not write “we used this for months”, “our team loved it” or “users find” without the evidence behind it.

Three states are kept apart for every run: whether the scenario was completed as specified, what the run observed, and whether the person who drove the session has confirmed the written observations. A completed scenario counts toward the testing status once its repeat pass succeeded; the driver’s confirmation gates statements of how the product felt to use, never statements of what was done. A scenario completed with a recorded workaround, such as a visit assigned to the account owner on a single-seat trial, counts as completed and is shown as such with the workaround named; the exception is recorded on the run, never applied silently, and a count is never loosened after the fact to keep a total. A scenario that was only inspected, with no repeat pass, does not count.

When a workflow test supports a published claim, the page cites it, and the citation resolves to a record like this:

Illustrative example, not a real test run
Hands-on verified

Example verification record

Product
Example product
Plan tested
Example plan
Workflow tested
Create an estimate, convert it to a job, invoice it
Test date
January 2026
Method
Hands-on, worked by a person in the product
Result
Workflow completed
Material limitation observed
The invoice step required a higher plan than the estimate step
Evidence label
Hands-on verified

How evidence becomes a recommendation

We do not publish a numeric score. A documented price or feature comparison can be stated as fact without any performance ranking; a comparative performance claim needs comparable hands-on evidence on each product for the claims that decide it. Recommendations are built from evidence and stated criteria, in this order:

  • Required capabilities are eligibility gates. A product missing a capability the page's workflow requires is not ranked for that workflow; it is listed with the gap named.
  • Completed comparable workflow evidence informs the assessment. What happened in a preserved test counts more than what a feature list says.
  • Price and buyer context can change the best fit. The plan a team size or a needed feature requires is part of the decision, not a tie-breaker.
  • Missing evidence is not estimated. A product without a completed test for a workflow is labelled "Not yet tested by us" for it, and no result is inferred to fill the gap.
  • Software-assisted observations have no independent evidentiary weight. Software may help with transcription, consistency checks, calculations and QA; it is never treated as hands-on experience, and it cannot create a subjective product claim.
  • A human reviewer approves subjective conclusions and final ordering. Timing is interpreted in context and is not treated as a universal speed score unless the test conditions make that comparison valid.

How products become eligible for a ranked order

  1. Eligibility. A product is considered for a page only when the evidence shows it can complete the workflow that page is about, for the buyer context the page is written for. A product missing a required capability is listed separately with the gap named, never quietly pushed down the list.
  2. Evidence threshold. A ranked order needs enough current, comparable evidence for the claims that decide it. Until that exists the page is a research list, not a ranking.
  3. Missing evidence. A product we have not run through the page's workflows is labelled "Not yet tested by us" for them: it carries no hands-on result, its recommendation status is limited to the evidence on record, and where a section is about a buyer context we have not evaluated it for, it is marked not evaluated for that context. No hands-on result is estimated to fill the gap.
  4. Buyer context. A different team size, trade need, integration requirement or budget can produce a different order, so the same product is considered from more than one buying position.
  5. Human approval. A ranked order is published only after the page meets its minimum evidence requirements and a human reviewer approves the conclusion. Before that point, products may appear in an unranked research list with their evidence and testing status, but not in a numbered or "best" order.
  6. Commercial independence. Commission, payouts, affiliate terms and any other commercial relationship are prohibited inputs to eligibility, evaluation, order and verdict.
  7. Exception disclosure. We do not make undisclosed ranking overrides. If a material methodology exception is ever required, it is documented on the affected page, cannot be based on commercial terms, and triggers a human review of the conclusion.

Ranking policy v2, adopted .

Incomplete evidence and retesting

Uncertainty is preserved on the page rather than designed out of it. These are the situations and what a reader sees in each:

Situation What the page says
A product has not been through our testing Its status reads “Not yet tested”. No hands-on result is assigned, and it is never ranked as though direct testing had occurred
Some of a product’s scenarios are tested and others are not “Testing in progress” with the count; the tested scenarios carry their findings, the others are labelled not yet tested
Testing ended before every scenario was completed, for example when a trial expired or a step needed a connection our rules do not allow “Core testing closed” with the count and the closing date; the evidence page names each scenario not completed and why; the completed findings stand with their dates, and no result is estimated for the rest
Vendor documentation contradicts what was observed The conflict is investigated. If it cannot be resolved and remains material, the page names the uncertainty instead of presenting either version as settled fact
A feature could not be verified directly It may still be described as vendor-documented or reported when those sources support it. If nothing supports it, the claim is not made
A price or plan limit may be stale The last verified date is shown; a stale fact is flagged for recheck and is not used for a current recommendation without that flag
The vendor prohibits automated access Automated access is not used. Permitted human testing is used where available; otherwise the limitation is disclosed
A conclusion needs a person’s judgment It is published as an editorial assessment, after human review of the relevant evidence, never as objective fact

Prices and plan limits are rechecked more often than workflow findings, because they change more often. Workflow evidence is retested after a material product change or when it goes stale, and a stale material fact is never used for a current recommendation without saying so. A material change to a conclusion carries a visible correction note with the date; the full process is in our editorial policy.

Affiliate independence

Commercial relationships do not determine which products are eligible, how they are evaluated, where they rank, or what verdict they receive. When a link can earn Service Tech Select a commission, the relationship is disclosed where the link appears. Products without an affiliate program remain eligible for recommendation, and a product that pays well is not recommended when the evidence says otherwise.

Service Tech Select does not currently use affiliate links. If that changes, a link that can earn us a commission will be labelled where it appears, and commission will never determine our rankings.

The affiliate disclosure lists any active program and explains how a commercial link is labelled.

Authorized testing, updates and policies

We use authorized accounts and permitted access methods, follow each vendor’s restrictions, and use synthetic rather than real customer data. We do not bypass logins, security checks, access limits or trial terms. Actions with real-world consequences require explicit human authorization. Where an account’s terms restrict publishing material from it, we describe what we observed rather than publish it, and retain the evidence privately for review.

Methodology updates

Material changes to how we evaluate software are versioned here. Results produced under different versions are not compared without accounting for the change.

Version Effective date Material change
v1.0 (clarification) Ranking rule update, same benchmark: a product we have not run through a page's workflows is labelled as untested and carries no hands-on result, and its place in a list comes from the documented evidence rather than being pushed below every tested product. Pages that recommend alternatives to a product now need the same human approval as rankings before a verdict appears. Nothing about scoring, scenarios or passes changed.
v1.0 (clarification) Clarification, same benchmark: when a vendor's own sources disagree, we name the controlling source where its scope and authority settle the question and explain the choice; where nothing settles it, the disputed detail is not presented as settled. Nothing about scoring, scenarios or passes changed.
v1.0 (clarification) Clarification, same benchmark: the testing status can read 'Core testing closed' when hands-on testing ends before every core scenario is complete, with the count, the closing date and the scenarios not completed named on the evidence page; a scenario completed with a recorded workaround, or inspected without a repeat pass, is shown as such. Nothing about scoring, scenarios or passes changed.
v1.0 Initial public methodology: the four evidence labels, human-run workflow testing with first-use and repeat passes, the core scenario suite and the chimney sweep trade module, ranking eligibility and human-review gates, affiliate independence, and the incomplete-evidence rules.
Every material change to how we test or rank adds a row here; results produced under different versions are not compared without accounting for the change. A clarification changes public wording or status vocabulary under the same version, so results stay comparable.

Corrections and policies

If something here is wrong or out of date, tell us. Corrections are the first priority in that inbox.