- The problem no one had formalized until now
- Problem architecture: how interference destroys rare memory
- SME use cases: when the model "forgets" what is really needed
- The solution: optimize your data before scaling up the model
- Trade-offs to consider before choosing
- What this study changes in how we evaluate models
- The recommended choice for Italian SMEs
A recent study has identified the precise mechanism that prevents small language models from acquiring rare skills. The problem isn't computational capacity in absolute terms. In fact, frequent tasks continuously overwrite what the model has learned on less common tasks. This phenomenon has been observed on models between 4 million and 4 billion parameters.
The most relevant discovery concerns the proposed solution. Instead of scaling the model to larger dimensions, it is sufficient to increase the frequency with which the target task appears in the training data. Therefore, SMEs considering the adoption of AI models do not necessarily have to opt for expensive enterprise solutions. A well-calibrated training data strategy can compensate for the dimensional difference.
We at SHM Studio we are monitoring this evolution carefully. The operational implications for Italian companies are concrete: choosing an LLM does not just mean comparing parameters, but understanding how it was trained and on what data. From this perspective, SHM Studio supports SMEs in evaluating and integrating AI solutions suited to their specific context, avoiding oversized investments compared to actual needs.
The problem no one had formalized until now
For years, the dominant narrative in the AI industry has supported a seemingly intuitive principle: bigger models yield better results. However, this statement hides an internal mechanism that until recently remained opaque. A new study, published and analyzed by The Decoder , has finally identified the precise mechanism behind this disparity.
Researchers analyzed models with a parameter range from 4 million to 4 billion. Within this range, they observed a systematic phenomenon. Frequent tasks in the training corpus continuously overwrite the learned representations for rare tasks. Consequently, small models don't fail due to a lack of absolute capacity, but because of a structural issue of interference between high and low-frequency signals.
This radically changes the perspective with which companies should evaluate language models. In fact, the question is no longer just ‘how many parameters does this model have?’. The correct question becomes: ‘what data was it trained on and with what frequency distribution?’.
Problem architecture: how interference destroys rare memory
To understand the mechanism, it's helpful to start with how an LLM learns during training. The model updates its weights at each iteration, trying to minimize the error across all tasks in the dataset. Therefore, tasks that appear more frequently generate stronger and more consistent gradients.
Rare tasks, on the other hand, produce sporadic updates. Every time a frequent task is processed, the weights shift in a direction that may be incompatible with what was previously learned on the rare task. This phenomenon is known in the literature as catastrophic forgetting , but the study in question detailed the dynamics in a more granular way.
In large models, this problem naturally diminishes. In fact, the greater parametric capacity allows for more stable representations to be allocated even for low-frequency tasks. However, the solution does not necessarily require increasing parameters. Increasing the frequency with which the target task appears in the training data produces a similar effect, at a significantly lower computational cost.
This distinction has direct implications for anyone designing fine-tuning pipelines on open source models or evaluating AI solutions for specific contexts. To dive deeper into the technical foundations of applied deep learning, MIT Technology Review offers an authoritative editorial perspective on these developments.
SME Use Cases: when the model "forgets" what it really needs
For an Italian SME operating in B2B or retail, this problem manifests in very concrete scenarios. Consider a company that uses an LLM to automate responses to support requests. Routine messages — requests for information on prices, hours, availability — are frequent, and the model handles them well. However, complex technical requests or structured complaints are handled inconsistently.
This isn't necessarily a problem with the model's intelligence. It's, most likely, a problem with the training data distribution. Complex tasks were underrepresented in the original corpus. Consequently, the model hasn't consolidated the necessary representations to tackle them reliably.
Similarly, a company using an LLM for SEO content generation might see excellent results for high-volume product categories and mediocre results for specific niches. Again, the likely cause is the frequency of exposure during training. Here at SHM Studio we regularly observe this pattern in the assessments we conduct for our clients.
For those managing integrated digital campaigns, the quality of AI output directly influences the performance of tools like google ads campaigns or the activities of SEO copywriting . Therefore, understanding the structural limitations of the chosen models is not an academic exercise, but an operational necessity.
The solution: optimize your data before scaling up the model
The study proposes a solution that is elegant in its simplicity. Before investing in larger models, it's worth checking if the problem can be solved by intervening in the training data distribution. In practice, this means increasing the frequency with which target tasks appear in the fine-tuning dataset.
This strategy has clear cost advantages. Large models require significant computational infrastructure, both for training and inference. Conversely, targeted fine-tuning on a compact model, with a suitably balanced dataset, can achieve comparable performance on specific tasks at a fraction of the cost.
However, this solution is not universal. There are tasks for which parametric capacity is genuinely necessary. Complex multi-step reasoning, handling very long contexts, and some forms of zero-shot generalization directly benefit from larger models. Therefore, the choice between a small optimized model and a large model remains dependent on the application context.
For SMEs, the actionable advice is to always start with an analysis of the distribution of real tasks that the model will face. This preliminary analysis allows you to correctly calibrate your training strategy and avoid oversized investments. Research from McKinsey confirm that most companies overestimate the complexity of the models needed for their actual use cases.
Trade-offs to consider before choosing
The choice between an optimized compact model and a large-scale model doesn't just boil down to performance. There are at least three trade-off dimensions that deserve attention.
- Inference cost: Large models require dedicated hardware or pay-as-you-go APIs with variable costs. Small models can run on-premise or on cost-effective cloud infrastructure.
- Latency: for real-time applications — chatbots, assistants integrated into e-commerce, sales support tools — response latency is critical. Compact models offer lower response times.
- Dataset maintenance: The data frequency optimization strategy requires continuous curation effort. This cost must be explicitly budgeted.
In addition to this, the dependence on third-party suppliers must be considered. Those who use proprietary model APIs have no control over the distribution of the original training data. In these cases, customization through fine-tuning or prompt engineering is the only lever available. To delve deeper into AI adoption strategies in business contexts, the SHM Studio AI services offer a structured starting point.
What this study changes in how we evaluate models
Before this research, evaluating an LLM for business use was mainly based on generic benchmarks. These benchmarks measure average performance across a wide set of tasks. However, for a company with specific use cases, average performance is a somewhat misleading metric.
What matters is performance on tasks that are actually relevant to the business. Therefore, the correct methodology involves building an internal benchmark, representative of real tasks, and evaluating models on that basis. Only in this way is it possible to identify whether the problem is parametric or if it can be solved through data optimization.
In summary, the study shifts the focus from model size to data quality and distribution. This is good news for SMEs, which rarely have the budget for enterprise models. It means that with a well-designed training data strategy, competitive results can be achieved even with accessible models.
For those managing activities of Digital marketing or SEO , this perspective opens up concrete scenarios for intelligent automation without the need for complex infrastructure. Our activities in web development already integrate this kind of logic into the design of AI-assisted interfaces.
The recommended choice for Italian SMEs
In light of what has been analyzed, the recommendation for an Italian SME evaluating the adoption or updating of LLM-based solutions is divided into three steps.
First, you need to accurately map the tasks the model will handle, distinguishing between frequent tasks and rare but critical ones. Then, check if the candidate models were trained on data distributions compatible with those tasks. Finally, before opting for large models, it's wise to test if targeted fine-tuning on a compact model, with a properly balanced dataset, yields sufficient results.
This approach allows costs to be controlled without sacrificing operational quality. For companies that want to delve deeper into these evaluations, the team at SHM Studio is available for a structured consultation. You can contact us via the page contacts or explore our Blog for further insights on AI and digital strategy.
For those who also manage activities on social platforms, it is worth considering how AI integrates with tools like the LinkedIn campaigns , where content personalization is a growing competitive factor.
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.