Skip to main content

The Questions AI Makes Possible

|

How LLMs are opening a new frontier for understanding how organizations really work

Futuristic data tree with glowing multi-color particles. Digital technology, information growth and analysis concept.
iStock/Hilch

In 2021, Ali Ahmadi, a doctoral candidate at Smith School of Business, decided to study something few management researchers had studied before: what happens when one company blindsides another. A price cut no one sees coming. An acquisition that catches a rival flat-footed. Ahmadi and his faculty advisor, Goçe Andrevski, were convinced the element of surprise had unique effects on competition that merited study. The problem was that no one had figured out how to measure the phenomenon. “In competitive dynamics, there was no real theory of surprise, no empirical studies of surprise,” he says.

The standard research method in strategy is to read news articles and manually log corporate actions — price changes, product launches, acquisitions, partnerships — then make a judgment call on which of those actions had startled the market. Ahmadi tried keyword searches, word counts and topic modelling, but none of these methods worked. Surprise, he realized, is found in context. It’s revealed in a journalist’s own words that suggest they, or the people they’re quoting, did not see something coming. “It’s impossible for traditional language processing methods to do that,” he says.

Then, in November 2022, ChatGPT arrived. Ahmadi didn’t trust it at first — “too weak and unreliable,” he thought at the time. But by 2023 he was ready to test it properly, feeding it articles pulled from the Dow Jones Factiva database. The results were good enough to prompt him to scale up using a larger database — two million news articles drawn from roughly 18 major outlets — that he processed through a large language model (LLM). For each article, the model identified which firm had taken an action, the kind of action, the target and, critically, whether the article described the action as surprising.

The result was a dataset of almost 6,000 surprising competitive actions, spanning more than 900 companies across 10 industries. It was a scale of measurement that simply did not exist when Ahmadi framed his research question three years earlier. “We couldn't even imagine one day we could do millions of articles manually,” he says, “but somehow it worked out.” His research question hadn’t changed. What changed was that it finally became answerable. 

Ahmadi’s breakthrough sounds like a story about research efficiency, but it’s the wrong frame. What he built wasn’t a faster research assistant but a new kind of instrument, one that will soon redefine what big management questions are even asked.

Bridging quantity and quality 

For decades, management research has run on two kinds of evidence. There is numeric data, in the form of stock prices, patent counts, executive compensation and survey scores, that’s clean, comparable and statistically tractable, but blind to anything that does not reduce to a number. And there is qualitative data — interviews, case studies, ethnographic fieldwork — that’s rich in nuance but limited by what a human being can read, code and compare by hand. A researcher could study six firms closely or 6,000 firms superficially, but not both. 

LLMs break the trade-off. These are AI systems capable of understanding and generating human language by processing vast amounts of text data. They can read earnings calls, annual reports, employee reviews, news coverage, court filings, even transcribed meetings and interviews at a volume no human team could approach, while still extracting contextual information that used to require a trained coder sitting with a highlighter. LLMs can infer tone, intent, relationships and meaning across millions of documents at once.

The latest LLMs are not even limited to text; they can now process speech, images and video alongside text. That opens possibilities for studying, say, leadership communication, meeting dynamics or emotional expression. 

The technology has triggered a wave of studies that would have been unthinkable or too labour-intensive even five years ago. In strategy, researchers are training LLMs to track how firms respond to rivals’ moves in close to real time, mining years of news coverage and case studies written after the fact.

Leadership scholars are parsing thousands of hours of meeting and earnings-call transcripts, performance reviews and other nontraditional sources of data for patterns in how leaders communicate, how CEO successions affect different dimensions of organizational culture and how different stakeholders perceive an organization’s culture — and where those perceptions diverge. 

In entrepreneurship, LLMs are being trained to classify startup narratives, evaluate business ideas, analyze crowdfunding campaigns and pitch language, and identify how ideas emerge and spread inside organizations.

Building the sim organization

Researchers are particularly excited to use LLMs to conduct “in silico” organizational experiments to test theories. Essentially, this involves creating virtual organizations staffed by managers (aka Homo silicus) with different incentives, organizational roles, experience levels or cultural backgrounds. In simulations, these actors interact, negotiate, compete and explain their decisions. Researchers can systematically manipulate one variable at a time — increase risk here, change incentives there — and observe how organizational dynamics emerge and how leaders make decisions.

Ahmadi’s own work illustrates where management is headed, and it goes beyond his construction of the unique “surprise” database. Not satisfied with cataloguing surprising actions, he pushed the LLMs further: What if you could capture every competitive action, by every firm, and understand how they related to one another over time? He trained an LLM model on almost six million articles, classifying each into a taxonomy of competitive moves while preserving the context around them — what technology was involved, where it happened and what strategic purpose it served.

The preserved context makes a second capability possible: using a function within LLMs called vector embedding to translate each sentence’s meaning into a set of numerical coordinates. This is a leap beyond older sentiment-analysis tools, which scored words in isolation and couldn’t distinguish a riverbank from a national bank. 

With each competitive action now carrying a numerical representation of its meaning, Ahmadi can compare actions across time and firms — and identify which move by a rival was actually a response to an earlier move, even if the two actions look nothing alike on the surface. And he could trace which of a firm’s scattered actions over months or years belong to the same underlying strategy — a pattern that would previously have required a researcher to read years of coverage and hold the entire narrative in their head at once. “It would take years to read all these articles and make sure we are part of the same grand strategy,” he says. “That’s what I’m doing right now.”

Ahmadi’s research is on the leading edge of management research, and this new direction should delight organizational leaders. Their companies contain much more strategic information than they can extract; the fact that academic researchers can begin to make meaning of all this unexplored data, and complete studies in a more timely fashion, could make management science more relevant to practitioners. It’s not unrealistic to imagine, for example, studies that connect recurring themes such as autonomy, psychological safety and recognition to productivity, turnover or innovation.

Mitigating the risks

None of this comes free of complications, and the researchers pushing hardest on these tools are also the ones most alert to their limits.

The most obvious risk is the one anyone who has used a chatbot already knows: LLMs can generate output that is confidently wrong. In a research context, that means a model might misread an article, misclassify an action or invent a relationship between two events that is not there. Ahmadi says the problem has diminished as the models have improved, but it hasn’t disappeared. As a check, after each processing run, he pulls a random sample of the model’s output and codes it himself, blind, then compares his judgments against the model’s classifications to check for false positives and false negatives. He also keeps his prompts narrow: a single, well-defined question per pass rather than a chain of reasoning steps, on the theory that every additional step in a model’s reasoning is another opportunity for it to drift. 

The ceiling imposed by the risks of LLM-driven research, however, is nothing like it was before. Since the early days of scientific management in the 1880s, a researcher who wanted to study something as elusive as competitive surprise or the emotional texture of a negotiation faced a hard choice: study it narrowly on a handful of firms or don’t study it at all. That choice is disappearing. Ahmadi’s own trajectory captures the shift. He set out in 2021 to answer a question that, by his own admission, was not really answerable, not with the tools available at the time. Today it is.

“We were looking at shadows on the wall,” Ahmadi says. “Now we can actually look at the actions first-hand.”