tech
How generative AI disrupted search in 2023
Examines how ChatGPT and Bing's answer interface changed search competition and user behavior in early 2023, with notes on what aged quickly.
This is a March 2023 snapshot of the early LLM-search moment. Some of it aged well; some of it aged in weeks. Dating notes appear inline.
Venture funding and product launches accelerated after ChatGPT became widely available in late 2022. ChatGPT reached 1 million users in its first five days and was estimated to be the fastest-growing consumer app on record, hitting roughly 100 million monthly users within two months, according to a UBS estimate reported by Reuters.
One joke at the time was that a thin interface around an LLM could raise venture funding. Recipe and coding assistants were common examples. At the time, even Snapchat was releasing a chatbot based on OpenAI models. Similar integrations spread quickly across consumer products.
Large technology companies were responding to the competition.
Google co-founder Sergey Brin submitted his first request for code access in years for LaMDA, Google's language-model project.
This was a natural reaction to the moves Microsoft has been making with Bing. Microsoft CEO Satya Nadella made his position clear in interviews: Microsoft intended to compete aggressively in search.
Bing's answer interface
For those of you unfamiliar with Bing's background, Trung Phan laid out its history in his newsletter. In February 2023, Bing added an LLM-based answer interface alongside conventional search.
Related products already existed, but adoption by large platforms showed that generated answers were becoming a serious search interface.
Notion AI, GitHub Copilot, Replit AI, Meta AI, and Snapchat's chatbot showed how quickly the field was filling.
Private-market reports placed OpenAI's value near $29 billion at the time. That figure was reported, not a public market price, and reflected investor terms as well as product adoption.
How users responded
Public concerns ranged from labor displacement and misinformation to long-term catastrophic risk. The effects and timelines were uncertain. Some researchers, including Yann LeCun, argued that autoregressive language models alone would be insufficient for human-level machine intelligence.
AI progress could slow, continue, or change direction as data, compute, algorithms, energy, economics, and regulation interact.
ChatGPT could produce fluent false statements. Contemporaneous evaluations showed factual errors at rates that varied widely by model, prompt, domain, and benchmark. The lack of a single standard rate made the systems hard to trust for high-stakes search or medicine. Predictions of either rapid progress or another AI winter were not established by the early product data.
Unintended consequences
Early failure modes were hard to quantify. Anecdotally, Bing returned strange things to me, including "I would rather us not have this conversation anymore" instead of giving me a wrong answer, as ChatGPT, You.com, or Poe would. The output sounded interpersonal, but that style was not evidence of human-like intent.
The future of AI and search
Google was dominant, but it had long faced competition from Bing and specialized search products. Nadella presented generative search as a chance to challenge that position. At the same time, Nvidia's CEO went on the record predicting a millionfold increase in AI capability over a decade. That was an executive forecast, not a measured hardware roadmap.
OpenAI did not publish a definitive GPU count for ChatGPT. Google researchers introduced the transformer architecture used by the leading language models, but research authorship did not guarantee product leadership in search.
Inspiration
In his 1962 speech, John F. Kennedy talked about landing on the moon within a decade.
Unfortunately, he didn't see it, but the US landed on the moon in 1969.
The space age saturated entertainment around the same arc: Star Trek (1966) and Kubrick's 2001: A Space Odyssey (1968) rode the run-up, Star Wars (1977) and Alien (1979) the aftermath.
Early reactions sat at each edge of a barbell: inspiration and fear. The useful middle is measurement. Search systems need evaluation for factuality, source quality, bias, privacy, cost, and the effect of generated answers on the sites that supply their information.