Machine learning is how software derives rules from past data and applies them to new data, instead of being programmed rule by rule. Business leaders do not need to understand the algorithms, but they need 7 concepts to evaluate an AI proposal properly: training data, models, predictions with probabilities, accuracy and how it is measured, overfitting, data drift, and large language models. With these 7 in hand you can tell which projects are feasible, what they cost to maintain, and which questions force a vendor to answer honestly.
Why leaders need this even with technical staff
AI investment decisions are usually made by non-technical people based on a vendor's presentation. Without foundational concepts you can only judge by feel, and feel is easily led by a polished demo.
The 7 concepts below are not for programming. They are the language for asking the right questions and reading reports after deployment.
1. Training data: the model only knows what it has seen
Machine learning learns from examples. To get a model that classifies warranty requests, you show it thousands of past requests that staff classified correctly. Three direct consequences:
- No data, no model. A new business with no recorded history cannot "use AI to forecast" anything yet.
- Bad data, bad model. If staff classified carelessly, the model learns that carelessness faithfully.
- Old data, stale model. A model trained on last year's customer behavior will be surprised by this year's.
So the first question for any AI proposal is: "Where does the training data come from, how much is there, who checked it?" From spreadsheets to smart systems is the data standardization step to do before any AI project.
2. The model: a tuned formula, not intelligence
A model is what comes out of learning: a mathematical function that takes inputs (customer information) and produces outputs (likelihood the customer will buy). It does not "understand" the customer. It found correlations in past data.
This matters because a model can be right for the wrong reason. A real example: a model predicting large orders learned that "orders placed at 9 a.m. on Mondays tend to be large", only because one big customer's purchasing department happened to order at that time. That customer left, and the model was completely wrong.
3. Predictions always come with probabilities
A model does not say "this customer will buy". It says "72 percent probability of buying". The business decides the action threshold: call customers from 60 percent up, or only from 80 percent?
A low threshold: many calls, more effort, more customers caught. A high threshold: fewer calls, some missed. There is no "correct" threshold; there is one that fits the business's costs and benefits. This is a business decision, not a technical one, and leaders should make it themselves.
4. Accuracy and how it is measured: 95 percent can be meaningless
"The model is 95 percent accurate" is the vendor's favorite line and the one leaders should doubt most. Example: 5 percent of transactions are fraudulent. A model that always says "not fraud" is 95 percent accurate and catches nothing.
You need two numbers instead of one:
- How many of the real cases it catches (what share of actual fraud the model finds).
- How many alerts are correct (of the cases the model flags as fraud, what share really are).
These two pull against each other: raising one lowers the other. Ask for both, and ask that they be measured on data not used for training, because on training data every model looks "great".
5. Overfitting: memorizing instead of learning
A complex model can memorize each training example instead of extracting rules, like a student who memorizes answers without understanding. The result: excellent scores on the old exam, poor scores on a new one.
The sign: test results look far better than real-world results. The prevention: always hold back a portion of data for checking, and run a trial on fresh data for a few weeks before trusting it.
6. Data drift: right today, wrong in six months
The world changes, the model does not, unless it is retrained. Prices shift, customer behavior shifts, new products launch, seasons differ. An inventory forecasting model trained on a year without promotions will be wrong in the first promotion month.
The budget consequence: an AI project does not end at handover. It needs continuous accuracy monitoring and periodic retraining. That belongs in the annual budget, not the one-off cost. Siri9 covers this in AI system management: weekly accuracy measurement, drift alerts, monthly updates.
7. Large language models: how they differ from "classic" machine learning
Everything above is classic machine learning: a dedicated model per problem, trained on the business's own data. Large language models (LLMs, the models behind ChatGPT, Claude, Gemini) differ in that:
- They are pre-trained on enormous amounts of text; the business does not train them but uses them through an API, paying by the amount of text processed.
- They do many tasks from instructions rather than training: reading, summarizing, classifying, writing, extracting information.
- They can fabricate fluently when they do not know. The common remedy is to have the model read the business's documents before answering, called RAG, explained in fine-tuning vs RAG.
For small businesses, most worthwhile work today uses LLMs through an API rather than training a custom model. Cheaper, faster, no large training dataset needed. But concepts 3, 4 and 6 still apply in full: outputs are probabilistic, must be measured properly, and drift over time (including when the provider changes the model version).
6 questions to ask before signing
- Where does the data for training or lookup come from, how much, who checked it?
- Which metric measures accuracy, on which data, and what are the two numbers from section 4?
- Who sets the action threshold, and can it be changed?
- What happens when the model is wrong, and who approves before anything reaches a customer?
- What is the monthly cost of monitoring and retraining, and who does it?
- Is our business data used to train anyone else's model?
A vendor who answers all six clearly has operated real models. One who answers "our AI is very smart" has not.
Where to start
Not with a model, but with one repetitive task that has data ready, done small, measured, with someone operating it. AI for small business, the 4 tasks to automate first lists those tasks and the selection criteria.
Frequently asked questions
Does a small business need to train its own model?
Almost never. Using an LLM through an API together with the business's documents solves most tasks at low cost. Custom training is only worth it with very large proprietary data and a highly specific problem.
Can "95 percent accuracy" be trusted?
Only when you know which metric and which data. Ask for the two numbers in section 4 and require measurement on data not used for training.
Can AI replace leadership decisions?
No. The model gives probabilities; people set thresholds and take responsibility. The leader's role is choosing tasks worth doing, setting thresholds, and demanding regular measurement reports.
How does Siri9 help at this stage?
A free assessment of an AI idea against the 6 questions above, the integration work under AI integration, and operations after handover. Send a description of the idea you are considering; we reply within 24 hours.
