Skip to content
Your cart

Your cart is empty. Let's fix that!

Search

Insights

Questions to Ask a Shopify Plus Agency in 2026

Questions to Ask a Shopify Plus Agency in 2026

The difference between a strong Shopify Plus agency and a weak one shows up in how they answer specific questions, not in how polished their presentation is. This is the full discovery-call script for a mid-market DTC brand hiring a Shopify Plus agency, organized into six categories: design fit, technical depth, integration capability, growth support after launch, process and communication, and commercial terms. Every question comes with answer grading so you can score responses during the call, not two days later from notes.

If you are still deciding how to evaluate agencies, start with the selection guide, which covers the five evaluation criteria, how to run a reference check, and what to budget. This post is for when the calls are booked.

Softlimit is a Shopify Premier Partner, a level Shopify awards for multi-million-dollar merchant impact and repeated Shopify Plus and Enterprise engagements. We have built on Shopify for 15 years. The questions below are the ones we would want a smart buyer to ask us.


Design fit: custom branded storefronts

The design section tells you whether an agency builds stores that express a brand or stores that wear a brand as a skin on top of a template. These questions are different from staffing questions. For design and development to do commercial work, the methodology has to start with the brand, not the build.

How do you translate an existing brand into a custom storefront, rather than adapting a template to fit it?

A strong answer describes a process: how the agency moves from brand assets (voice, visual language, purchase psychology, core customer) to a design system built for the store from the beginning. It should name specific brand attributes from a past client and explain how those shaped layout and interaction decisions, not just color and type.

A weak answer is a portfolio tour. Showing you builds that look good and offering to do something similar is not a methodology. It is a gallery.

Can you show a build where a specific design decision moved a commerce metric?

A strong answer names the store, the decision, and the result, and connects them. At Softlimit, engineering a custom product detail page to manage 23,000 variants within a single seamless experience for The Perfect Jean was a design-and-architecture decision, not a visual one. Order volume rose 200%. That is not a result to claim but a mechanism to explain.

"When we launched The Perfect Jean, we knew we didn't want a cheap build or an order-taker. We wanted a true thought partner. Softlimit understood that immediately. They brought deep Shopify expertise, fashion experience, and a collaborative mindset that let us build the site right the first time."

-- Ovadia, Co-Founder, The Perfect Jean

A weak answer is screenshots and awards. Proof is a named result with a reason. A render is not proof.

Who designs and who builds, and what is the handoff between them?

A strong answer names a process: how a design becomes a specification, how that spec travels from design to development without losing fidelity, and who is responsible for holding both sides to it. Joint review sessions, shared component libraries, and a defined handoff moment are real mechanisms. Ask which one they use and what breaks when it does not work.

A weak answer is "we have a design department." Separate departments without a defined interface produce drift between what was designed and what ships.


Technical depth

These questions tell you whether the agency builds for how Shopify works or around it. The distinction matters because over-engineered builds cost more to maintain, more to update, and typically require a rebuild sooner than they should.

On your last build, which Shopify-native capabilities did you use where a third-party app was the obvious shortcut?

A strong answer is specific: Functions, B2B, Markets, Checkout Extensibility, or another native surface, with the reason that capability was the right call for that client's situation. The agency should be able to explain what the app alternative would have cost the client in ongoing maintenance and update fragility.

A weak answer is an app list. Knowing which apps exist is not the same as knowing when not to use them.

How do you handle a catalog with deep variant architecture?

A strong answer describes the product data model, how filtering stays within Shopify's limits without breaking the product detail page experience, and where custom logic is genuinely required versus where native merchandising handles it. A 23,000-variant catalog for The Perfect Jean is what this problem looks like at scale: engineered to work as a single seamless PDP, not split across multiple workarounds.

A weak answer is "it depends on the catalog." Without elaboration, that is not an architecture answer. It is a delay.

Walk me through your QA process before launch. What does it actually cover?

A strong answer names the elements: responsive breakpoints at real device sizes, edge-case data scenarios (products without images, empty cart states, zero-inventory products), accessibility checks against WCAG 2.1 AA, cross-browser functionality, and the staging-to-live promotion sequence. A real process has a checklist. Ask for it.

A weak answer is "we test everything before launch." That is an intention, not a process.


Integration capability

Most Shopify Plus builds fail at the integration layer, not the design layer. These questions tell you whether an agency thinks about your stack as a system or as a feature list.

Before recommending a new tool for our stack, what do you audit in the existing one?

A strong answer starts with an inventory: what the current stack does, what it costs, and where tools overlap or conflict before any new tool enters the conversation. The right first question is whether the current setup already does it well enough. An agency that leads with its preferred tool list before understanding yours is replacing your stack with its defaults, not solving your problem.

A weak answer is asking which apps you currently use without a framework for evaluating them.

Tell me about a subscription or ERP integration you have shipped. What made the handoff between platforms reliable?

A strong answer names the client, the platform, and the specific mechanics: where the integration logic lives, how payment tokens transfer, what the cutover sequence looks like, and who is responsible when something goes wrong between systems. Softlimit migrated Moon Juice to Recharge in 25 days against a hard contract deadline, with subscription continuity protected throughout, because the cutover was sequenced around revenue protection before any data moved.

A weak answer is "we have done several" with no mechanics. Anyone can claim several. What matters is whether the agency can describe how.

If we remove an app the build depends on, what happens to our data?

A strong answer addresses data portability, export formats, and whether customizations are app-dependent or exist in the theme code and documentation the client owns. It distinguishes between what you own and what lives in a SaaS subscription the agency controls.

A weak answer is "that would be a conversation with the app vendor." If the agency built around an app, it should understand what removing it involves. If it does not, that is a build risk, not just an exit risk.

For the broader framework on evaluating integration decisions before the build, see our integrations overview.


Growth support after launch

Post-launch work is where most agencies disappoint and where a mid-market build's commercial value is actually realized. These questions tell you whether an agency has a growth model or a support ticket queue.

What does the work actually look like in months two through six?

A strong answer describes a cadence: CRO hypotheses defined before launch, specific areas of the funnel being measured and iterated, and a clear distinction between maintaining the build and developing it. Months two through six are where growth happens or does not. "We are available for whatever you need" describes reactive support, not a plan.

A weak answer is availability without structure.

How do you measure and report on post-launch performance?

A strong answer names the metrics tracked (conversion rate, average order value, page load time, revenue per visitor), the reporting cadence, and the format. A shared dashboard both parties can read, not a monthly PDF that arrives without context, is the right standard. Strong answers connect the metrics to the business goal the build was designed to move. At Softlimit, Verb relaunched in 8 weeks with Rebuy integrated into the experience before launch; Rebuy drove a 59% ROI and a 15% lift in AOV within the first 30 days, a result published in Rebuy's own case study. That result started with a hypothesis set before the build, not after it.

A weak answer is "we send a monthly report." Without knowing what it contains or how it drives the next sprint, a report is documentation, not accountability.

What exactly does the retainer buy, and what falls outside it?

A strong answer names what is included (development hours, strategy, CRO iteration, priority incident response) and is equally clear about what triggers a separate scope conversation. The clearer the boundary, the more reliably both parties can plan. Strategy and growth retainers should be structured as a growth engagement, not a support ticket queue repackaged as partnership.

A weak answer is a vague hours bucket. "25 hours a month for whatever you need" without prioritization logic is maintenance, and calling it a growth retainer does not make it one.


Process and communication

These questions reveal whether an agency runs a build or improvises one. The structural questions belong here; the staffing and boundary questions (named project lead, subcontracting, who you are not right for) live in the selection guide.

What is the communication cadence during the build, and who owns the project channel?

A strong answer describes a specific rhythm: standing calls at a defined frequency, async updates in a named tool, and one person who owns the project communication channel and is responsible for it staying current. The structure matters more than the platform.

A weak answer is "you will have full access to the team." Access to a team without a structure is noise.

How are scope changes handled once a build is in flight?

A strong answer describes a formal process: the change is documented, the impact on timeline and cost is estimated before the work starts, and the client approves before scope expands. Agencies that have run this process recently can describe a specific case. Ask for one.

A weak answer is "we are flexible." Flexibility without a process means scope expands and the deadline holds until it does not. Then neither holds.

What happens if something breaks after launch? Who classifies the severity, and how fast does it move?

A strong answer names a severity classification, a response time by severity level, and a clear escalation path. The strongest answers also address what happens during Shopify platform incidents, not just code bugs, because not everything that breaks in your store broke in your code.

A weak answer is "we will be there for you." That is a sentiment. Ask for the service level behind it.


Commercial terms

Commercial terms questions rarely get asked directly enough. An agency comfortable with its pricing and contract structure will answer these plainly.

How is pricing structured, and what specifically triggers a change order?

A strong answer describes the pricing model clearly (fixed-fee with defined scope, time-and-materials with a cap, or another structure) and names what triggers a change request: scope added after discovery, integration complexity not visible until kickoff, client-side delays that compress timelines, third-party dependencies that shift. The agencies that have shipped builds at your complexity level know exactly where change orders come from. They say so.

A weak answer is "we will handle it," which means either the agency absorbs scope creep until it stops, or the client absorbs it without warning.

At Softlimit, retainers start at $5,000 per month on a 25-hour floor, and builds anchor from the mid five figures into six figures by complexity. Scope drives where any given project lands within that range.

What is explicitly excluded from your quote?

A strong answer names the exclusions before you ask: third-party app licensing, Shopify subscription cost, content production, photography, copywriting, custom app development beyond a defined boundary, and anything dependent on third-party APIs outside the agency's control. Exclusions named in advance are not surprises.

A weak answer is a quote that does not mention exclusions. The exclusions are still there. They are just waiting.

What is the contract length, and what does exit look like?

A strong answer describes the notice period plainly and whether any tooling creates an ongoing dependency on the agency. An agency confident in its work does not need contractual lock-in to retain clients. The handover-ownership question itself (theme code, configurations, documentation) belongs to the selection guide.

A weak answer is multi-year terms with ambiguous exit language. That language was written by someone who expected you to need it.


How to use this checklist

Bring it to every discovery call. Grade each answer in the meeting, not later. Compare agencies on what they actually said, not on how the presentation felt.

For questions about staffing, who owns your project, subcontracting, handover asset ownership, and what kind of client an agency is not right for, those questions belong to the selection guide. This checklist and that guide are designed as a pair: the guide tells you what to look for; this post tells you what to ask and how to grade the answer.


Frequently asked questions

What should you ask a Shopify Plus agency in the first call?

Cover six areas: design fit (how they translate a brand into a custom storefront rather than a template), technical depth (which native capabilities they use and why), integration capability (how they audit an existing stack before adding to it), growth support (what post-launch development actually looks like), process (how scope changes are managed mid-build), and commercial terms (what the quote includes and excludes). You do not have to get through all 18 questions in one call. One question per category in 30 minutes will tell you more than a portfolio review.

How do you tell a strong agency answer from a weak one?

Strong answers are specific, name a number, and explain the connection. "We engineered a 23,000-variant PDP and order volume rose 200%" is specific, named, and connected. "Our clients love what we build" is none of those things. The test is whether the answer is falsifiable: if you cannot verify it against a named project, it is positioning, not proof. Strong agencies know the difference.

How many agencies should a mid-market brand interview?

Three is a workable number. Fewer risks a false comparison. More diffuses the signal. With a structured checklist, the gap between a strong and a weak answer is usually clear by the second call. What takes longer is finding three agencies worth calling. That is where the shortlist work matters more than the number of calls you book.


Softlimit is a Shopify Premier Partner with 15 years on Shopify and the results to show for it. Run the checklist on us. If the questions above are the right ones to ask, let us answer them.

Let's Talk Shop(ify).