Do I Need a Data Scientist to Do This?
For most first AI deployments at mid-market scale, you don't need a data scientist. You need a different kind of expertise. Knowing which is which prevents expensive hiring decisions that don't solve the actual problem.
- basics
- ai fundamentals
- hiring
- implementation
One of the most common wrong moves I see mid-market companies make when starting an AI program is this: someone decides they need a data scientist. The job description goes up. A hire happens. Six months later the company has a highly credentialed person with nothing useful to work on, because the actual first-year problems didn't require a data scientist at all.
The role confusion is understandable. AI sounds technical. Data science sounds like the technical discipline most adjacent to it. But for the large majority of first deployments at mid-market scale, these are different skill sets pointing at different problems. Getting clear on which is which saves you from an expensive mismatch.
Platform agents need a platform specialist
The most common first AI deployment at a company this size is deploying an AI agent inside a platform the company already runs. Salesforce Agentforce. HubSpot Breeze. ServiceNow. Microsoft Copilot Studio inside the 365 environment. These platforms have AI capabilities built in, and activating them is a configuration and implementation problem, not a research or modeling problem.
What you need for this: a platform implementation specialist. Someone who knows the specific platform architecture, knows how agents are configured within it, knows the relevant data model and integration points. This person looks more like an advanced CRM or ERP consultant than a data scientist. They're working within a defined system, configuring behavior, and connecting it to your existing data. Data science skills don't help here. Platform expertise does.
Custom automations need a developer
The second common first deployment is a custom workflow automation calling an LLM API. Your company has a workflow, an intake process, a document, a stream of data. You want to build an automation that calls GPT-4, Claude, or Gemini to process it. The output feeds into your existing systems.
What you need for this: a software developer with AI integration experience, plus someone who can design the workflow architecture. The developer needs to know how to call the API, handle rate limits and errors, manage prompt versions, and wire outputs into downstream systems. The workflow designer needs to know which steps to automate, how to handle exceptions, and how to define what "correct behavior" looks like so you can test for it. This is software engineering and workflow design. Still not data science.
Where a data scientist actually belongs
Custom model training is where data science expertise actually appears in the requirements. If your organization has proprietary structured data, and you want to build a model specifically trained on that data to predict something, you need a data scientist to design and maintain that model. That's a real use case.
Almost no mid-market company should be doing this for a first deployment. The pre-trained models available today, GPT-4, Claude, Gemini, and their variants, handle the vast majority of mid-market use cases without custom model training. They're trained on more data than your company will ever generate, across more domains than your internal team could specialize in. Custom training requires a clean, labeled dataset (which most companies don't have), engineering infrastructure to run training jobs (which most companies haven't built), and ongoing maintenance to keep the model current (which is expensive). The typical ROI calculation doesn't work at first-deployment scale.
The intelligence layer, where pattern recognition across your specific operational data becomes genuinely predictive, is year-two or year-three work. It requires a functioning data pipeline built on the records created by the automations you run in year one. Before those records exist, there's nothing for a data scientist to model. Building this capability before you've built the data foundation is a common and expensive sequencing error.
Ask which of the three you're doing
For a company starting its first AI deployment, the honest question isn't "should we hire a data scientist?" It's "which of these three things are we actually doing: configuring an agent in an existing platform, building a custom workflow automation, or doing something more sophisticated?" For the first, hire or contract a platform specialist. For the second, hire or contract a developer with AI integration experience, or engage an implementation agency that brings this capability embedded in the project. For the third, you probably aren't there yet, and the agency worth working with will tell you that directly.
Before you hire anyone, ask an implementation agency what role they would fill in your first engagement. A credible agency will describe the specific expertise the project requires, not give you a generic answer. If the answer is "a data scientist," ask them specifically what modeling work they're proposing and why a pre-trained model won't do the job. That question tells you a lot about whether they understand your situation or are proposing what they have available.
If you want that question answered for your specific situation, the Forge Playbook does it. Answer a few questions about your business and we'll put together a tailored outline of which workflows are worth automating and what a realistic budget looks like for each. Free, no obligation, takes about three minutes.