Development is no stranger to hype cycles. Since the turn of the century, the field has undergone needless cycles of over-promise and under-delivery with, just to name a few, the randomistas, blockchain, mobile connectivity, and microfinance. The field’s reception of the newest kid on the block, artificial intelligence, shows that lessons from the past still bear repeating.
In the 2010s, increasingly cheap phones and widening connectivity meant more services would reach the last mile. The poorest would gain access to services they had never before been able to reach. More than fifty years since the world’s first mobile phone call, income-based disparities in phone ownership persist. The smartphone is still a luxury good owned by only around one in four of the poorest sub-Saharan Africans.

Those who do have phones are still hindered from making full use of the internet’s vast knowledge. Web search promised to give people full access to the world’s knowledge. However, use of the internet remains lower to mobile penetration in many low-resource countries, mainly due to the cost and availability of bandwidth.1
At the end of the 2010s and early 2020s, it was now the blockchain’s turn. Tech leaders and development professionals alike became increasingly excited. Blockchain technology was now being deployed for payment and identification systems for refugees, land registries, and other applications. As a case study of a humanitarian cash-transfer pilot in Jordan illustrates, some pilots were blockchain in name only (a “glorified spreadsheet”, according to some), a sign that the usefulness of presenting interventions as using blockchain to stand out in the philanthropic market.
Relative to their promise and buzz, the solutions of the day have often left more muted, but still significant, marks in the field. The randomista revolution did not solve the problem of evidence, but it did push towards more thoughtful use of fit-for-purpose evaluations, contributing a huge body of knowledge on what works in different fields. Tech has also delivered: M-Pesa, for instance, allowed Kenya to massively expand financial inclusion without relying on large financial institutions and bank accounts.2
These processes of over-promise and under-delivery often follow the same patterns of over-excitement and over-commitment. Gains, if they happen, tend to be smaller and take more time than anticipated. Sometimes, they do not materialize at all.
Pilots, but make it AI
With great hype comes what some in the development community call “pilotitis”, the uncoordinated proliferation of proof-of-concept pilots that, regardless of results, do not scale. This is perhaps the biggest shame of the hype-cycle pattern development is stuck in, signifying vast amounts of wasted opportunity.
In the mobile health craze, this largely led to a graveyard of pilots. The situation was so bad that, in 2012, the Ugandan government put a moratorium on mHealth pilots to try to get a handle on the number of NGO requests for partnership in testing a large number of overlapping (but not interoperable) mHealth solutions. With no real chance of scaling, most of these pilots did not lead to large-scale deployment or adoption.

It is now artificial intelligence’s turn on the bandwagon. It is not rare to hear that advances in robotics and artificial general intelligence will soon make most of the problems that plague low-resource settings seem like molehills.
In the end, low income countries’ problems are due to a variety of structural factors, all conspiring against easy progress. These are physical and political bottlenecks, which require costly commitments to be overcome. The development community, well aware of these challenges, should not fall victim to a collective dream that imagines a world where digital tools can overcome the circumstances that have kept countries poor for centuries.
Short-termist gains
Yet attention is elsewhere. As Iqbal Dhaliwal, J-PAL’s executive director, warns: development loves a silver bullet. There’s also pressure to satisfy the interests of donors who get excited by newness and innovation. New programs and incubators promise funding and access to compute credits for narrow applications, such as weather forecasting and clinical decision support. While laudable and—potentially—worthwhile, these do not address the fundamental transformation challenge posed by AI.
These priorities also fail to address the current political context. Growing resentment against AI and anxiety for what the future may hold are manifesting in politically salient ways, such as backlash against data center investments in the US and tensions over sovereignty and access to tools built by others. For low- and middle-income countries, it materializes in having to choose between different provider partnerships, and wrestle with how much control over citizen data they wish to hand over.
Perhaps because transformational goals look unattainable (and expensive), development practitioners and governments have an incentive to prioritize short-term narrow use cases of AI-tools. Low-resource settings may not be ready to leverage AI to the fullest, but that isn’t stopping government agencies, the private sector, and non-profits from moving ahead with adoption.
For instance, around 60% of NGO-led community-health-worker (CHW) programs surveyed by the Community Health Impact Coalition now use an AI tool, yet only one of 28 programs reported meaningful integration with the public health system. That’s unsurprising: two-thirds of respondent NGOs said the systems they operate in weren’t ready for AI.
The mismatch between AI adoption at the programmatic level and the lack of preparedness in the overall system raises concerns about the (lack of) effective governance of AI in low-resource countries, and the dim prospects for scale and sustainability.
It is not all gloom—AI does have a role in making services cheaper and more effective. The UK government, for instance, has export promotion teams that use machine learning to identify export-ready businesses, bill managers use LLMs to predict parliamentary outcomes, and local councils are testing AI-based flash flood warnings for farmers.
Even so, hype outpaces rigor. In an update to the CHW survey mentioned above, only a few of the 44 programs documented were being rigorously evaluated. The survey also found that design details and basic facts such as models used under the hood are hard to pin down, and claims of effectiveness and scale are difficult to validate or inconsistent. Perhaps unsurprisingly, there’s much similarity to a 2018 review of blockchain use-cases in international development that “found no documentation or evidence of the results blockchain was purported to have achieved in these claims” across 43 use cases.3
The use of AI is now taken as synonymous with impact. A report from GSMA typifies the sector: AI “drives”, “powers”, “enables”, “generates”, and “integrates”—but high-quality evidence that it achieves its aims is scarce. GiveDirectly, for instance, has shown how they use models to serve their core purpose by helping with flood prediction (independently verified) and documenting how users interact with AI services to inform design.
Evaluating AI tools will not require randomized trials for every application, with practical solutions such as measuring intermediate outcomes or A/B testing potentially filling gaps in the evidence. At the moment, the evidence we create is usually difficult to generalize due to model performance drift and variation in how tools are used across contexts. It also lags because model improvements outpace the academic publishing timelines.
The case for a smart bets agenda in AI for Global Health and Development
Amid rapid and uncritical adoption, it is worth taking a step back and avoiding the mistakes of previous hype cycles in development. At the macro level, we need a politically aware and feasible set of priorities for how low resource countries can start to improve their preparedness for AI. And in a narrower sense, we need more thoughtful evaluation of the AI use cases that can most cost-effectively support people at scale.
These questions are hard enough to answer for established fields in Global Health and Development, let alone fast-moving innovations like AI.
The smart buy
From the name, one can picture a harried policy analyst who—on demand from a superior—rushes to the policy shop and navigates the aisles searching for the right policy proposal. Facing the familiar choice paralysis, the analyst lays eyes on the World Bank endorsed, perfectly placed, “smart buy” product. Trusting the five star review, the analyst heads to the ministry, smart buy in hand.
Reality is not so simple, as critics of the smart buy have pointed out. Analysts rarely have free rein—they are constrained by implementation capacity, what will allow for political gains, and what already fits into preconceived notions of what works.
But people do seem to pay attention. At their best, “smart buy” or “what works” efforts in different fields can provide a strong signal for the types of things that tend to work best, and what experts who are leaders in their field understand to be the evidence base.
The smart buy agenda traces back to evidence-based medicine. Disease Control Priorities (DCP1) was a companion to the World Bank’s 1993 World Development Report “Investing in Health”. DCP1 provided the basis for the Bank’s argument that maximizing health drives development. Its cost-effectiveness approach using disability-adjusted life years is reputed to have influenced Bill and Melinda Gates to focus on health.
In the development space, the “smart buy” has become almost synonymous with the Global Education Evidence Advisory Panel (GEEAP), which has released two reports in 2020 and 2023, recommending cost-effective approaches to learning in low- and middle-income countries. Convened by the Foreign, Commonwealth and Development Office (FCDO), the World Bank, UNICEF, Rachel Glennerster, and Abhijit Banerjee, the project was bound to make noise.
Rigorous evaluation of these publications’ impact is limited, but they appear to have boosted structured pedagogy and teaching-at-the-right-level, the approaches they most champion. FCDO and World Bank project documentation shows a marked increase in mentions of both approaches in recent years. Pinpointing best buys in different sectors was central to Rachel Glennerster’s strategy at FCDO and remains part of the department’s strategy of influence on other actors.
Others have followed suit, including the WHO’s best buys on non-communicable diseases and non-profit reviews on violence against women and the Copenhagen Consensus ‘best things first’ efforts.
At their best, “what works” agendas align funders, development finance institutions, and practitioners.
The concept of “smart buys” is not without its detractors.
In 2004, famed development economist Lant Pritchett was hired to write the education paper for the original Copenhagen Consensus, which was among the first identifiable efforts to make concrete general recommendations of the “smart buy” sort in the development field.
His task was to write a “challenge” paper that would, among other things, list cost-effective education interventions. Editors instead received a paper detailing “reasons why ‘recommendations’ about how to improve education had to be based on a correct positive model of what education producers were actually doing and why”—and refusing to produce a list. The organizers pointed to the terms of reference; Pritchett offered to return his fee (the paper was accepted in the end, and makes for excellent reading).
Pritchett’s refusal of a specific list of recommendations makes the task of clean advice difficult (especially if you are his editor and you have pre-committed to publishing such a list). However, policy failures are not usually a matter of lacking technical knowledge of what is the best option, but about the system actors and their incentives. Pressure to make specific interventional recommendations detached from specific context risks overstating how valid the evidence is.
A sketch for a future AIDev smart bets panel
Detractor concerns are worth taking seriously, but building an advice agenda for AI in development would look quite different from previous efforts. For one, there is little evidence to speak of, so advice will likely have to take the shape of early principles and guidance on smart evaluation. Additionally, the speed of what is technically possible risks making advice stale fairly quickly, meaning that the effort is likely to be most useful if it takes a dedicated, always-on, approach.
Previous smart buy efforts have usually taken the shape of an academic exercise: systematic evidence reviews carried out by researchers and presented to a reputable board of advisors who meet and shape the evidence into a package of recommendations. In some cases, like the GEEAP, the evidence base evaluation is maintained on an ongoing basis and the recommendations updated every couple of years.
This process of systematic review paired with expert discussion has worked well. A more skeptical take is that beyond the discussion and research, the value comes from achieving prominence and influence through excellent communication, and the persuasion of influential donors. Getting donors on board, and aligning their approaches towards asking for the same things, is immensely powerful.
For the AI use case, rather than packaged solutions, the advice needed is more akin to good principles for investment. This panel could employ researchers to maintain a public and live literature review of the ever-evolving academic and gray literature on both narrow implementation trials and the potential macro-economic effects of AI on low- and middle-income countries. The panel, convening respected and prominent implementers, policymakers, academics, and individuals with strong technical understanding of the underlying technology, would come up with principles for investment which would hopefully have staying power among the development community.
The benefits of such a set up are that it could coordinate industry, governments, and the academy around new technological developments in artificial intelligence, while hopefully acting as a source of restraint against the tendency to proliferate pilots without an overarching theory of change. The panel may very well recommend rather mundane investments—the boring stuff that makes leveraging AI possible: connectivity, digital infrastructure, and workforce adaptation, to name a few. These topics have a longer history and evidence base, and advice given may prove useful under a multitude of future scenarios.
The group could also partner with those building fit-for-purpose evaluation frameworks. Evaluating ever-evolving technological capabilities requires a different approach. Newly released guides from J-PAL and a joint effort by the Center for Global Development, The Agency Fund, and IDinsight are strong initial efforts in the right direction.
Previous smart buy efforts have been shown to be powerful rhetorical tools that brought attention and coordinated actors, usually in decades-old fields. A panel working on such a new issue may prove premature. However, it is when noise is at its highest, signals are pointing in different directions, and hype is pushing dumb investment that this type of effort is most needed.
An AIDev panel of respected and independent individuals, equipped with the best evidence we have and their own experience, may be just what we need to start cutting through the noise.
This is shown rather neatly in this pre-print by Björkegren et al. that identifies cost-efficiencies and increased use for a mobile messaging based AI chatbot, relative to web search, for Sierra Leonean teachers.
M-Pesa’s success stands as an example of smart investment by development agencies working with private companies. The initial investment for M-Pesa came from a £1 million pound prize by the United Kingdom’s Department for International Development (now Foreign, Commonwealth and Development Office) matched by £1 million by Vodafone. Stemming from the observation that Kenyans were at the time already trading airtime, mobile money then grew to be commonplace worldwide.
This is not a problem in development alone. A 2024 study published in JAMA reviewed over 900 FDA-approved AI-enabled medical devices: barely half had clinical performance studies and only 2% had randomized trials. The authors found FDA information “frequently insufficient” to assess clinical generalizability.
