Automotive consumer research should break down “which car do you like” into six types of evidence: candidate set formation, information comprehension, family negotiation, dealership visits and test drives, purchase and delivery, and long-term usage.Conceptual preference only indicates choices under given materials and conditions; purchase intention is not an order, and high ratings after a test drive only represent those who reached the test-drive stage. Research that genuinely supports product or channel decisions must clarify which stage consumers are in, who participates in the decision, what the actual usage conditions are, and what behaviors will be used to further validate the conclusions.
Automotive decisions involve high stakes, long cycles, and are influenced by price changes, financing options, used-car disposal, delivery timing, energy refueling, living conditions, and family members. Putting everyone into a single “brand and configuration preference” questionnaire tends to yield seemingly stable but non-actionable average answers. Research should focus on one decision at a time—for example, determining which range explanation is easier to understand, why appointments after test-drive requests did not result in dealership visits, or which scenarios within the first three months after delivery did not match pre-purchase expectations.
Define the Research Population by Stage and Role First
- Need Emergence: Clear trigger to replace or add a vehicle but no specific model list yet; suitable for studying tasks, budget boundaries, and alternative transportation arrangements.
- Candidate and Comparison: Already actively searching or comparing models; suitable for studying brand entry points, information sources, energy options, and trade-off rules.
- Dealership Visit and Test Drive: Already scheduled, visited, or completed a test drive; suitable for studying real tasks, sales interactions, and candidate set changes.
- Order and Delivery: Already placed an order, cancelled, or completed delivery; suitable for studying contracts, fees, waiting periods, expectations, and vehicle condition.
- Usage and Replacement: Sufficient real-world ownership experience; suitable for comparing scenario performance, refueling/charging, software, after-sales, and next purchase considerations.
Within the same household, the primary driver, co-users, purchaser, and information gatherer may not be the same person. Role-type questions are insufficient here; instead, ask what each person actually did: who raised the need, who screened candidates, who test-drove, who paid, and who will use the vehicle long-term. For joint household decisions, interview separately and then compare; do not let one respondent guess the genuine concerns of others.
One Thousand Lead Cases: Forty Percent Conversion Is Actually Only Nine Point Six Percent
Below is an illustrative funnel. Starting with one thousand people who meet the definition and have recent purchase plans, six hundred and twenty requested model brochures, three hundred scheduled test drives, two hundred and forty completed the test drive, and ninety-six ultimately placed an order. If the dealership simply divides orders by the number of completed test drives, the calculation is:
96 ÷ 240 = 40.0%
This figure can describe the “order rate among those who completed a test drive,” but it cannot be called conversion among all those planning to purchase. With the initial one thousand as the denominator, the order rate is only 9.6%. The show-up rate from scheduled to actual test drive is 80%, and the rate of scheduling after requesting information is approximately 48.4%. The three denominators correspond to three different decisions: improving candidate entry, reducing no-shows after scheduling, or improving post-test-drive closure.
If questionnaires are only sent to the 240 test-drivers, the team will never see the reasons from the 620 who did not request information, the 320 who viewed information but did not schedule, or the 60 who scheduled but did not show. No matter how high the test-driver ratings, they cannot prove overall market acceptance. Recruit separately from each stage and preserve the source, model, region, time period, and stage disposition codes.
Real-World Usage Conditions Before Configuration Voting
Questions like “Which is most important: range, space, power, or driver assistance features?” force consumers to choose among abstract words. A more effective approach is to record one week of typical trips, the longest common one-way distance, the number of passengers, cargo carried, parking and refueling/charging conditions, winter and summer climates, highway trip frequency, and alternate vehicle availability—then let consumers make trade-offs among clearly priced configurations and option packages.
Energy costs should not be reduced to a single national average either. The U.S. EPA states that standardized fuel economy estimates are for model comparisons and do not guarantee every driver achieves the same real-world result; driving behavior, roads, climate, and vehicle maintenance all cause variation. Research can use uniform assumptions for fair comparisons while asking participants to recalculate scenarios with their own mileage, energy prices, and refueling structure—clearly distinguishing standardized values from personal assumptions.
Four Hundred-Person Randomized Test: Comprehension Improved, Consideration Showed No Significant Change
A team tested a revised EV range communication. The current version highlighted a single standardized range figure; the new version retained the same figure but clarified that it is suitable for cross-model comparisons, actual performance varies with speed, temperature, road, load, and air conditioning use, and provided three typical scenario range explanations. The model, price, visual order, and other parameters were identical. Four hundred qualified near-term car buyers who had not seen either version were randomly assigned to two groups of two hundred each.
The primary comprehension metric was whether respondents correctly judged that “the standardized figure is for consistent comparisons across models; individual results may vary.” In the current version, seventy-eight answered correctly (39%); in the new version, one hundred and thirty-two answered correctly (66%), a difference of twenty-seven percentage points. Using a simple approximation for two independent proportions, the 95% confidence interval is approximately 17.6 to 36.4 percentage points. The evidence supports that the new version improved comprehension.
As for consideration into the candidate set, one hundred and twelve in the current version selected consideration (56%) versus one hundred and eight in the new version (54%). The difference is minus two percentage points, with an approximate interval of minus 11.8 to plus 7.8—so it cannot be concluded that the new version changed consideration. A qualified conclusion is that “after providing fuller explanations of risks and usage conditions, comprehension improved and no clear change in consideration was observed,” rather than “transparent explanations have no effect on sales” or “consideration dropped by two points.”
This remains a material test, not a road performance validation or a sales forecast. Next steps could include observing the new version in real channels regarding information search, test-drive questions, configuration choices, and post-delivery expectation gaps, while maintaining pricing and promotion records. If market conditions, model availability, or incentives change later, order fluctuations cannot be entirely attributed to the wording.
Test-Drive Research Should Record Test Conditions, Not Just Driving Impressions
Before the test drive, record the candidate set, expectations, and the most important tasks to verify; immediately after the drive, note whether tasks were completed, where sales explanations were needed, and which expectations were confirmed or refuted; a few days later, track changes in the candidate set and next-step behaviors. Route, duration, traffic, weather, vehicle configuration, state of charge or fuel, occupants, sales accompaniment, and waiting times must all be logged; otherwise, differences across dealerships may simply reflect differences in test-drive conditions.
T?tasks should come from real-world scenarios, such as installing a child seat, loading typical luggage, navigating tight parking spaces, understanding dashboard prompts, or locating commonly used controls—rather than vague “feel the handling.” Extreme tasks that cannot be safely completed during a normal test drive should not be forcibly simulated for research. Questionnaires can record understanding and confidence, while engineering, safety, and compliance conclusions must rely on appropriate professional testing.
Driver Assistance Research Must First Standardize Terminology and Responsibility Boundaries
Consumers may conflate manufacturer feature names with driver assistance and autonomous driving categories. The U.S. NHTSA notes that differing manufacturer naming causes comprehension difficulties and distinguishes between warnings, emergency interventions, and steering or acceleration/braking assistance; its consumer-facing information emphasizes that current assistance levels still require sustained driver engagement and attention. Specific capabilities are subject to model, market, version, usage conditions, and official documentation.
Therefore, research cannot simply ask “do you trust intelligent driving?” First present accurate and equivalent feature descriptions, then test capability boundaries, driver responsibility, activation conditions, exit prompts, and failure comprehension separately. Self-reports of “very safe” or “fully hands-off confidence” are not evidence of safety performance; rather, they may signal communication or mental-model risks. Feedback involving accidents, malfunctions, or potential harm must enter formal safety processes, not remain as open-ended topic themes in a report.
Order, Cancellation, and Delivery Require Event-Level Denominators
Reasons for order cancellation may include price changes, delivery delays, financing rejection, family opinion shifts, competitor launches, or configuration misunderstandings. Survey first establishes the order timeline, promised delivery window, actual notifications, change events, and the sequence of cancellation, then asks why. Asking customers to multi-select from a fixed list of reasons collapses the event process into post-hoc attribution.
Delivery experience also distinguishes among vehicle condition, documentation and fee comprehension, feature walkthroughs, waiting periods, and problem resolution. Satisfaction on delivery day does not guarantee usage fit three months later; early novelty, seasonality, and software versions can alter the experience. Long-term tracking should preserve model, batch, software version, region, climate, and typical usage while managing repeat contact burden.
Dealership, Community, and Owner Samples Each Miss Different Groups
Dealership samples cover only those who visit, brand owner communities skew toward highly engaged owners, after-sales samples skew toward those who require service, and online panel samples rely on screening accuracy. Research can combine sources but must tag source, report unweighted base sizes for each source, and compare observable structures. Competitor owners, no-show visitors, cancelled orders, and former owners cannot be represented by brand-active owners.
Data such as income, financial details, precise location, VIN, or full account identifiers should only be collected when decisions genuinely require it and after ethical review; use valid ranges and regional levels where possible. Combining home address with fixed parking and charging conditions may increase identification risk and should not be imported simply for analytical convenience. Publishing small-sample regions or rare model combinations also requires protecting individual identity.
Frame Findings as Verifiable Product or Channel Actions
Recommendations should clarify the level of intervention. If those who did not schedule a test drive were unaware of the relationship between standardized range and personal scenarios, the action is to modify comparison explanations and test comprehension; if no-shows after scheduling concentrate on weekday evenings with transportation barriers, the action may be adjusting times or test-drive methods; if post-test-drive rejection stems from child-seat installation difficulty, it is a space and task issue. These three situations cannot all be written up as “strengthen sales training.”
A decision table should include at least: research population and denominator, observable facts, consumer interpretations, competing hypotheses, controllable actions, validation metrics, and safety guardrails. When behavioral facts align with interview interpretations, the evidence is stronger; when they conflict, preserve the tension. For example, if consumers say range is most important but consistently choose lower-priced options in real selections, the interaction of budget constraints, social desirability bias in questioning, or configuration combinations may be at play—do not cherry-pick only the evidence that supports an existing approach.
Pre-Publication Checklist
- The target population is defined by purchase stage, model or energy type, and planning timeframe.
- Roles of primary driver, co-users, purchaser, and information gatherer are distinguishable.
- Each of preference, consideration, scheduling, test drive, order, delivery, and usage uses its correct denominator.
- Configuration comparisons include real prices, packages, and usage conditions—no abstract feature voting.
- Standardized energy consumption or range is clearly separated from individual scenarios, without promising individual results.
- For test drives, preserve route, model, version, weather, accompaniment, and task completion conditions.
- Driver assistance terminology, capabilities, and responsibility boundaries are standardized first; self-reports do not substitute for safety validation.
- Dealership, community, owner, and panel samples are each described with their sources and coverage gaps.
- Experiments pre-specify comprehension, behavior, long-term outcomes, and safety guardrails.
- Reports distinguish observable facts, consumer interpretations, research inferences, and causal evidence.
For deeper dives, purchase stages and touchpoints can be expanded with customer journey research; concept and configuration comparisons may reference product concept testing; sample coverage and extrapolation limits relate to target population definition; and randomized material tests follow randomization and order control. The most important aspect of automotive research is not obtaining a high intention rate, but knowing where that rate sits in the funnel, under what conditions it was derived, and how many unvalidated steps remain before real orders and long-term usage.
Automotive Research Sources
- U.S. EPA: Transportation, Air Pollution and Climate Change FAQs: Standardized fuel economy estimates are for comparisons; actual results vary with driving, roads, climate, and vehicle conditions.
- U.S. EPA: Text Version of the Fuel Economy Label: Explanation of energy consumption, annual costs, and uniform assumptions.
- U.S. NHTSA: Driver Assistance Technologies: Driver assistance feature classification, naming differences across manufacturers, and the boundary of sustained driver engagement.
- AAPOR: Best Practices for Survey Research: Requirements for population, sampling, recruitment, questionnaires, mode, weighting, and public disclosure.

