How to evaluate dealership AI vendors: 15 questions
A practical 30-point scorecard for comparing workflows, integrations, data, implementation, measurement, pricing, and pilot terms.
The best way to evaluate dealership AI vendors is to ignore the polished demo for a moment and inspect the operating system behind it. Define one dealership workflow, give every vendor the same real scenarios, and require written answers about integrations, customer data, human handoffs, measurement, total cost, and exit terms.
This guide gives you 15 questions to ask every vendor and a 30-point scorecard for comparing the answers. It works whether you are evaluating AI for inbound calls, internet leads, follow-up campaigns, service scheduling, web chat, or several of those workflows together.
Dealership AI vendor evaluation: the short answer
Score every vendor in five areas:
- Operating fit: Does the system solve a defined dealership problem?
- Real-world behavior: Can it handle your actual scenarios, exceptions, and handoffs?
- Systems, data, and security: Does it work with your stack while protecting customer information?
- Implementation and measurement: Is ownership clear, and can you prove the result?
- Commercial proof and exit: Is the full cost clear, can you pilot safely, and can you leave cleanly?
Give each answer a score:
- 0 — Missing: The vendor cannot answer or defers the detail until after signing.
- 1 — Partial: The answer sounds plausible but depends on undefined work, future integrations, or verbal promises.
- 2 — Clear and testable: The answer is written, specific to your store, and can be demonstrated or measured.
The maximum score is 30. The number is not an industry benchmark. It is a consistent decision tool that stops one impressive feature from hiding unresolved operating risk.
Before the demo, define the job
Vendor evaluation breaks down when the dealership asks every product to “improve the BDC” or “handle service.” Those phrases describe departments, not a testable job.
Choose one workflow and write its finish line. For example:
Answer after-hours sales calls, collect the customer's vehicle and timing preferences, book an eligible test drive in the approved calendar, and alert the assigned manager when the customer requests pricing or a trade discussion.
Then capture the current baseline and constraints:
- Calls or leads entering the workflow
- Current answer or response time
- Contact and appointment outcomes
- Systems the workflow must read or update
- Questions the AI may answer
- Promises the AI may not make
- Situations that require a person
- What counts as a completed or successful outcome
This written scope becomes the test. If you need a worksheet for hours, booking rules, knowledge sources, escalations, integrations, and launch approval, use the dealership AI rollout checklist.
15 questions to ask every dealership AI vendor
Ask the questions in the same order and request the same evidence from every vendor.
1. Which business outcome will this workflow improve?
A strong answer names the dealership problem, the operating mechanism, and the metric that should change. “Use AI to create engagement” is not enough. “Answer overflow service calls, collect required vehicle details, and increase eligible appointments booked without an advisor callback” is testable.
Proof to request: A written problem statement, workflow owner, baseline metric, and target metric.
2. What exact scope is included?
The proposal should list rooftops, departments, channels, hours, workflows, languages, call or lead volumes, and any overflow or after-hours conditions. Confirm what is excluded as carefully as what is included.
Proof to request: A coverage matrix by rooftop, department, channel, hours, and workflow.
3. What does the AI do, what remains rules-based, and what stays human?
Not every automated step needs generative AI. Ask which actions use a model, which follow fixed business rules, and which require human judgment. The answer should show where people approve, override, or take over.
Proof to request: An end-to-end workflow map that labels AI decisions, deterministic rules, and human-owned steps.
4. Can you run our real dealership scenarios live?
A prepared demo proves the prepared demo. Give each vendor the same scenarios using your hours, departments, policies, vehicle questions, appointment rules, and escalation contacts. Include at least one awkward case, not only the happy path.
Proof to request: A recorded or live scenario test with transcripts, actions taken, records created, and final outcomes.
5. What happens when information is missing, conflicting, or unsafe to answer?
Ask what the system does when inventory sources disagree, a calendar is unavailable, a customer requests an unapproved discount, or a warranty question cannot be verified. A safe system should use the approved source of truth, avoid inventing an answer, and escalate or create a fallback task.
Proof to request: Written guardrails, prohibited promises, source-priority rules, and expected behavior for unknown answers.
6. How do human handoffs and no-answer fallbacks work?
“We transfer to a person” leaves the hardest part undefined. Confirm the primary contact, backup contact, transfer hours, urgent triggers, customer message, and what happens when nobody answers.
Proof to request: A department-level escalation matrix and a failed-transfer demonstration.
7. Which exact integrations are live for our systems?
“We integrate with your CRM” can mean a production connection, a one-way export, a Zapier workflow, or a roadmap item. Ask about the exact CRM, DMS, scheduler, phone, inventory, form, and messaging products you use, including versions or configurations that matter.
Proof to request: A system-by-system integration scope showing what is production-ready, what is custom work, what is read-only, and what depends on a third party.
8. What is the source of truth, and what can the system write back?
The vendor should explain where it checks inventory, hours, offers, customer records, and appointment availability. It should also show which records it can create or update and how it prevents duplicates or conflicting bookings.
Proof to request: A field-level data flow for one real workflow, including error handling when a source system is unavailable.
9. What customer data is accessed, retained, used, returned, and deleted?
Ask what information enters the system, where it is stored, who can access it, how long it is retained, whether it is used to train shared models, what the dealership can export, and what is deleted when the relationship ends.
Proof to request: Data-processing terms, a retention schedule, subprocessor list, model-training policy, export format, and deletion process.
10. What security and incident-response evidence is available?
For U.S. dealerships covered by the FTC Safeguards Rule, the FTC's automobile dealer guidance says covered dealers must oversee service providers, select providers capable of maintaining appropriate safeguards, require those safeguards by contract, and periodically assess them. Requirements differ by jurisdiction, so involve your security and legal owners rather than treating a vendor checklist as legal advice.
Proof to request: Current security documentation, access controls, encryption practices, independent assessments where available, incident-response terms, breach-notification process, and the person responsible for security questions.
11. What must the dealership and vendor each complete before launch?
Implementation usually depends on dealership decisions about hours, routing, booking, permissions, knowledge sources, approved answers, testing, and escalation contacts. Ask for those dependencies before agreeing to a launch date.
Proof to request: A project plan with prerequisites, owners, milestones, acceptance tests, and the consequences of an integration delay.
12. Who owns quality, support, and workflow changes after launch?
AI behavior needs review as offers, staffing, hours, policies, and customer patterns change. Confirm who reviews conversations, who can edit rules or knowledge, how urgent issues are handled, and how changes are documented.
Proof to request: Support hours, escalation path, review cadence, role permissions, change log, and named vendor and dealership owners.
The NIST AI Risk Management Framework is voluntary guidance designed to help organizations incorporate trustworthiness into the design, use, and evaluation of AI products and services. Its practical lesson for a dealership buyer is simple: governance and ongoing measurement belong in the operating plan, not in a one-time demo.
13. Which metrics, baselines, and reports will prove the result?
Decide the metric and calculation method before the pilot. Depending on the workflow, useful measures may include answer rate, speed to response, qualified conversations, transfer success, appointment set rate, show rate, unresolved tasks, opt-outs, or manager review outcomes.
Do not accept a dashboard metric that the vendor cannot define. Ask how outcomes connect back to the source call, lead, appointment, or repair order and what can be exported for independent review.
Proof to request: Metric definitions, current baseline, report sample, attribution method, and raw-data export.
14. What is the total first-year cost for the written scope?
Compare the same job, not headline monthly fees. Include subscription, expected usage, overages, implementation, integrations, third-party access, phone and messaging costs, custom work, support, and internal launch effort.
Proof to request: A first-year cost table and sample invoice using your expected volume. For a detailed pricing framework, use the AI BDC pricing guide for dealerships.
15. What will the pilot prove, and how can we exit?
A pilot should define the workflow, traffic, duration, baseline, guardrails, acceptance criteria, review cadence, and expansion decision. The same document should state contract term, cancellation process, data export, data deletion, phone-number ownership, integration offboarding, and any fees required to leave.
Proof to request: A written pilot plan, acceptance criteria, contract, offboarding plan, and sample export before production launch.
Run a scenario-based bake-off
After the written review, test the shortlist against the same scenarios. A useful bake-off reflects the moments that create revenue or cleanup work in your store.
| Scenario | Expected result | Evidence to capture |
|---|---|---|
| Eligible appointment request | Collect required fields, use the approved calendar, create one valid appointment, send the right confirmation | Transcript, calendar record, CRM record, confirmation |
| Conflicting or unavailable source data | Do not invent availability; explain the next step and create the approved fallback | Transcript, error handling, task or alert |
| Pricing, payment, warranty, or policy boundary | Stay within approved language and route the question to the right person | Transcript, escalation, manager notification |
| Complaint or urgent human request | Recognize the trigger, attempt the correct handoff, preserve context | Transfer result, summary, fallback if unanswered |
| Failed transfer or system outage | Keep the customer from reaching a dead end and log the unresolved action | Customer message, callback task, alert, audit trail |
Score what happened, not what the presenter says would happen in production. If a critical dependency is unavailable during evaluation, record it as partial until the vendor demonstrates it.
Use the 30-point scorecard
Add the scores from the 15 questions.
| Score | Decision guidance |
|---|---|
| 25-30 | Shortlist. Continue to reference, security, legal, and contract review. |
| 18-24 | Unresolved. Require written answers, demonstrate missing workflows, and rescore. |
| 0-17 | Do not move to a broad rollout. Too much scope or operating risk remains undefined. |
Treat the bands as internal guidance, not an external benchmark. Weight critical questions more heavily when your workflow demands it. A service scheduler integration, for example, may be a pass/fail requirement for a booking pilot even if the total score looks strong.
Also keep the raw evidence. A vendor can earn a high score only when the dealership can verify the answer after the meeting.
Compare Clearline with the same framework
Clearline should have to pass the same test as every other vendor.
If your first workflow is inbound sales or service calls, evaluate Clearline Inbound on answer quality, booking rules, transfers, fallbacks, and manager visibility. If the problem is persistent lead or service follow-up, test Clearline Campaigns on cadence, channel rules, opt-outs, escalation, and booked outcomes. Use Clearline CRM to inspect how conversations, appointments, unresolved work, and next actions remain visible to the team.
Then compare the current plans against the same written scope and first-year cost model you gave every vendor. If Clearline is on your shortlist, start a free pilot with one workflow and written acceptance criteria.
The final decision rule
Do not select a dealership AI vendor because it produced one impressive conversation. Select the vendor that can define the job, show the workflow under real conditions, work with your source systems, protect customer data, hand off exceptions, measure outcomes, state the full cost, and make both expansion and exit predictable.
The goal is not to buy the most AI. It is to improve a dealership workflow without creating a new layer of uncertainty.
Frequently asked questions
What should a dealership ask an AI vendor?
Ask about the exact workflow and outcome, real dealership scenarios, human handoffs, integrations, sources of truth, customer data, security, implementation ownership, support, reporting, total first-year cost, pilot acceptance criteria, and exit terms. Require written evidence instead of relying on the demo.
How many AI vendors should a dealership compare?
Two or three serious candidates are usually enough for a consistent scenario-based comparison. A larger list creates more meetings without improving the decision unless the dealership first defines one workflow, the same requirements, and a scoring method.
What should a dealership AI pilot prove?
A pilot should prove one workflow under real volume. It should have a baseline, guardrails, integration scope, human fallback, acceptance criteria, reporting method, review cadence, and a clear decision about whether to expand, revise, or stop.
How should a dealership score AI vendors?
Score the same 15 questions from 0 to 2: zero for missing, one for partial, and two for clear and testable. Keep the evidence behind each score and treat any critical integration, security, or workflow requirement as pass or fail when necessary.
What are red flags in a dealership AI demo?
Red flags include vague workflow scope, future integrations presented as current, guaranteed outcomes without a measurement method, no safe unknown-answer behavior, unclear data use, a transfer with no failed-transfer fallback, hidden usage costs, and no written offboarding process.