Can AI Actually Do Innovation Discovery? What the Evidence Shows, and What a System Has to Do About It
AI fails where it is asked to decide, not where it is asked to generate. That distinction is what makes a trustworthy innovation discovery system possible.
September 1, 2026
Ask an operations leader what is on their improvement list, and you get a number. Thirty items. Sixty. A hundred and change, spread across four spreadsheets and two people's heads. Ask which one to fund first, and the answer changes shape. It stops being a list and becomes a story about who asked loudest, what broke most recently, and which number someone felt able to defend in front of the board.
The methods for doing the work are mature. Lean, Six Sigma, constraint analysis, stage-gate. All of it exists; all of it works when it is pointed at the right thing. The unsolved part is upstream of the methods. It is knowing what to point them at, and being able to say why that one and not the other twenty-nine.
That is the job AI ought to be good at. Putting far more candidates on the table than any team has time to generate, drawn from your own operations, adjacent industries and problems that look nothing like yours until they do, then handing back a ranked set you can take to leadership and defend.
And the first thing most people say to that is: but it makes things up.
They are right, and the research backs them. What the research also shows, once you read it closely, is that the failures aren't scattered randomly. They cluster somewhere specific, for a reason you can name, and once you can name it, you can design around it.
So the short answer is yes, but not on its own. AI is strong at producing breadth: generating candidates, framings and angles, and making sense of material put in front of it. It is weak at deciding what is true or what ranks highest. Searching is a separate matter, and the record below is unkind to a model left to run its own search unsupervised. A discovery system you can trust keeps the model on the first job and moves the judging into fixed code, where human judgment is written down once and applied the same way every time.
What is AI actually good at in innovation work?
The evidence here is not thin. It is uneven in a way that proves useful.
Inside what the model handles reliably, the gains are real and large. Dell'Acqua and colleagues, publishing in Organization Science in 2026, ran 758 early-career consultants at Boston Consulting Group through 18 realistic product-innovation subtasks in a randomized field experiment. The group with access to GPT-4 completed about 12% more tasks, finished roughly a quarter faster, and produced work rated between 1.3 and 1.5 points higher on a ten-point quality scale. That is one global professional-services firm on the April 2023 model, and the authors make it clear that effects may differ elsewhere.
AI generates volume, and the volume is decent. In a circular-economy business-idea challenge, Boussioux and colleagues compared 125 eligible human submissions against 180 AI-generated ones across 3,900 evaluator ratings. The AI ideas scored slightly higher on average value and slightly lower on novelty. They were also far more similar to each other than to anything a person had written. The apparent edge in the best ideas was too small to separate from chance, and the study measured what evaluators thought, not what happened when anyone tried to build one.
It widens what a person notices. A controlled lab experiment with 124 entrepreneurs found that AI assistance increased the number of opportunities identified, and reduced novelty and sensitivity to context in the same sample (Cristofaro, Giardino and Muldoon, Technology in Society, 2026). That is one lab study rather than a field result, and it is close to the whole of the credible base on this question.
Read those together, and a pattern shows up. Every one of those wins is a generation win.
So why does AI get things wrong so confidently?
Because of what it is, and it is worth understanding rather than working around.
Most people expect software to behave like a calculator: same inputs, same answer, every time, and you can check the work. A language model is not that. It is a probabilistic system. It samples, word by word, from a range of plausible continuations, leaning toward the likely ones without being bound to them. Ask it the same question tomorrow, and you may get a slightly different answer, because there was never only one answer sitting in there to begin with.
Everything else follows from that. A model is fluent whether it is right, because fluency and correctness come off the same production line. Nothing inside it separates a number it read in a source from a number that merely fits the shape of the sentence. Sounding sure and being right are not different operations in there, and the finished text carries no mark to tell them apart, unless something outside the model supplies one.
Which means the model has nothing to check itself against. It is not that it grades its own work badly; there is no separate faculty in there doing the grading.
That is the whole failure, and you can see it in that same consulting-firm experiment. On one task in that experiment, the same population was 19% less likely to be correct, and the work they produced beyond that boundary was rated more coherent than the control group's, not less. The output got worse while the writing got better. The researchers call the boundary a jagged technological frontier, and its defining property is not that it exists. Every technology has limits. It is that the boundary is invisible from where you are standing when you cross it.
One further detail changes what you would do about it. The consultants who had been given a prompt-engineering overview did worse past the frontier than the consultants given the model alone. That comparison was observed inside the study rather than designed as a test of prompt training, so treat it as an observation. But it is the only evidence available on whether better instructions fix the problem, and it points the wrong way.
Better prompting is not the answer. The answer has to sit outside the model.
Where does the evidence run out?
We went looking for research on the discovery-stage work that matters most: catching weak signals early, challenging what a team already believes, and ranking opportunities against each other.
On assumption challenge, AI as a devil's advocate against the room's consensus, we found nothing. No weak study, no mixed result. Across a full research pass, nothing peer-reviewed or independently verifiable. The only adjacent material was vendor marketing.
On prioritization and scoring, the same. The entire evidence base reduces to one consulting-firm blog post with no sample, no study, and no tool evaluation.
Weak-signal detection has five distinct, verifiable studies behind it, and none report how often the method was right. All five demonstrate a method rather than measure it: Mühlroth, Kölbl and Grottke on separating innovation signals from noise (Scientometrics, 2023); Ha, Yang and Hong on keyword-network clustering with a graph convolutional network (Futures, 2023); Miao, Guo and Yuan on identifying AI-industry directions from weak signals (IEEE Transactions on Engineering Management, 2024); Lee, Kwon, Kim and Kwon on early identification of emerging technologies from patent indicators (Technological Forecasting and Social Change, 2018); and KISTI's institutional self-report of 439 signals across 24 fields in 2023. The single named business outcome is one retail anecdote, relayed second-hand, with no disclosed method and no control.
One strong negative result is worth sitting with. Clark and colleagues (2025), a systematic review of 19 studies drawn from 3,071 screened records in Research Synthesis Methods, found that when generative AI was used to search the literature unsupervised, it missed 68% to 96% of relevant studies, and wrongly excluded studies that belonged between 1% and 83% of the time. The review concludes that current evidence does not support using generative AI in evidence synthesis without human oversight. That work sits in clinical and library-science methodology, so read it as a parallel to business research rather than a measurement of it. But the shape of the failure is the same everywhere else: it fell apart on the decisions (what to include, what to leave out) while running unsupervised.
So: five things people want AI to do at the front end of innovation. One has rigorous evidence, and that evidence says do not run it unsupervised. One has a thin, mixed base. Two have nothing at all.
That silence is not a gap in our reading. It is the record's current state, and it is the most interesting fact in this whole piece.
It is also not slowing anyone down. Deloitte's 2026 Global Human Capital Trends surveyed more than 3,000 business and human-resources leaders across 15 countries and every industry: 60% of executives now regularly use AI to support their decisions. Asked about addressing what that does to decision rights and accountability, 64% called it important to current success, and 62% had efforts underway, while 5% reported making great progress. Deloitte's own forecast is worth quoting exactly: as organizations expand AI-enabled decision-making, many "find AI to be amplifying existing deficiencies instead of solving them."
What does that tell you about how to build the system?
Everything from here is our position. It is reasoned from the failures above, not tested by them, and we will not put a citation next to a claim no study supports.
Line up the failures, and the pattern is consistent. The wins are generation. The failures are judgment. The model is strong when it is asked to produce possibilities, and unreliable when it is asked to decide: what is true, what counts, what ranks above what.
That is not a limitation to prompt your way past. It is an instruction about where the model belongs.
So we stopped asking how to make AI more reliable at deciding, and asked a different question: what if it never decides anything? Let the model do what it is demonstrably good at: generate breadth, and make sense of what is put in front of it. Then take the act of judging away from it. That does not remove judgment from the system. It moves the judgment into the open: written once by people, held in fixed code, applied the same way on every run, and available to be inspected and argued with. The final call stays with the person accountable for it.
Say that out loud, and it sounds almost too simple. It is not simple to build.
What does an innovation discovery system actually have to do?
Here is the honest anatomy, at the level of what each part is for rather than how it is built. There are roughly eight jobs. The model does two of them.
1. Intake. Capture this specific business: what it does, what it can do, what it cannot, and the constraints it actually operates under. Without this, the system answers a question about your industry instead of about you.
2. Aiming. Anchor the search to a fixed frame of where value gets created in a business of that shape, and to the company's own operations. This is what stops a wide search from becoming a wide search over the wrong ground. Our categories for the forms innovation can take are fixed and defined in advance, grounded in established discovery methods rather than invented on the fly, so the model works from settled terminology instead of making up a category and then reasoning on top of it.
3. Generating breadth. (This is where the model works.) Be exact about the division of labour here: retrieval tools go and fetch the real sources, because the record is clear that a model left to run its own search unsupervised misses a great deal. The model frames what to look for and works with what comes back. Deliberately many angles, many framings, many adjacent domains. Breadth has to be engineered. One broad prompt is not a search, and a model asked one question one way returns the sensible middle of its distribution. This is where the probabilistic behavior is the asset rather than the liability: it will raise the option nobody in the room would have written down.
4. Synthesizing. (This is also the model.) Turn what was found into candidate opportunities a person can actually evaluate.
5. Carrying the evidence. Every material number arrives attached to where it came from, all the way through to the finished page. Nothing becomes true just because it sounds true.
6. Grading, from outside. How strong a claim is gets computed in code, from the quality of the sources behind it. The model never reports its own confidence. This is the direct answer to a model having nothing to check itself against.
7. Disqualifying and ranking. Fixed rules that throw options out, then a fixed formula over a fixed set of dimensions. No model anywhere in the arithmetic. Same inputs, same ranking, every time, which means you can inspect it, argue with it, and re-run it.
8. Checking itself. When a line of research is genuinely thin, the system flags it rather than generating support that is not there. Consistency gets measured across runs, not assumed.
Six of those eight exist for one reason: to make the two the model does trustworthy enough to act on. That is the part the market conversation keeps missing. The sophistication is not in the prompt.
Does a system like this need to know your industry?
This is the skeptic's real objection, and the answer is more interesting than yes or no.
No. It does not need to know your industry the way a consultant who has spent twenty years in it knows your industry. What it needs is the shape: the handful of stages where work flows through a business of that kind, and the activities that support them. That is a small, fixed frame, and it is what aims the search.
The depth comes from you. Your capabilities, your constraints, your operating reality, supplied at the start. So the honest response to "AI doesn't know my business" is: correct, and it is not supposed to. You point it at your business. It searches. Fixed rules score what it brings back.
That is also why the same system flexes across very different industries without being rebuilt for each one. We run against six industry shapes today, at different depths, and the runs show the steering works: the frame changes, the machinery does not.
The number our own system invented
We reached all of this the slow way, on our own product.
An early Hephanos run produced a customer-facing figure: 60% referral leakage. We stated it as present-tense fact. We threaded it consistently through eight or more report fields. It was directionally plausible, because referral leakage is a real phenomenon at roughly that magnitude in that industry. The system labelled its own confidence in that number as high.
It had no source. The research behind it carried no citations at all. And that confidence label was not a judgment about evidence quality. It was computed from how many supporting points the model wrote. Four or more, and the system called it high. The model had been asked to grade its own homework with a rubric that measured answer length.
Nothing in the output announced the problem. That is the entire difficulty with this technology in one example. A confabulated number does not arrive flagged. It arrives fluent, consistent, and reasonable, which is exactly the profile of a number worth trusting.
Our fix was not a better prompt. We do not believe a better prompt was available. We moved the grading out of the model.
What we have not solved
We are pre-customer. We have no case studies, and this honest version of the argument doesn't pretend otherwise.
What we do have is volume. The system has run hundreds of times, and we have reviewed the analysis behind those runs in depth, not one report at a time, but across runs, looking for where results drift, where similar inputs land differently, and which outputs sit outside the pattern. That is not a customer outcome. It is the evidence available before customers exist, and it is what the design decisions above are calibrated against.
It has not produced a perfect system, and a discovery system tuned for perfection would be the wrong trade. The line between what is plausible and what is defensible is thin, and a system that never crosses it will also never surface the opportunity that only looked implausible at first. Now that generating candidates is cheap, the binding constraint is no longer cost. It is the quality of what makes it through. Putting the ranking in fixed code is how you take that trade and still defend what survives.
Some parts are further along than others. A few of the checks we have built run as warnings rather than hard stops today. They mark a weak number and flag the report's quality rather than holding it back, and moving them to hard stops is work we have authorized and not finished. The step that scores raw material before the ranking runs is the least settled part of the system; we measure how consistently it orders the same inputs, it is not yet where we want it, and we track that gap rather than describe it as closed.
What would turn any of this from position into evidence is founding-cohort results with the method disclosed. Those do not exist yet. We are building toward the claim, not standing on it.
Why not build this yourself?
Almost all of it is fixed cost. The defined categories, the disqualifying rules, the plumbing that carries a citation from a source all the way to a finished page, the self-checks, the deterministic scoring core and the calibration behind it. None of that gets cheaper because you only ever point it at one company. And none of it is finished work. Every piece has to be maintained against models that change underneath it.
Spread across many companies, that cost is ordinary. Rebuilt inside one, it becomes a standing engineering commitment in service of a system that is not what your business is for.
To be clear about what is shared: not ideas. Every run produces its own set. The intake differs, the constraints differ, the research differs, and the generation is probabilistic, so no two companies receive versions of the same answer. What travels between companies is the machinery, never the output.
The one question to ask any vendor
Two systems can put an identically shaped ranked list on your screen and be built on opposite premises. The screen does not tell you which one you are looking at.
One question does, and you should ask it of anyone in this category, including us: which parts of your system are probabilistic, which parts are fixed rules, and why is each one where it is?
Ours is the eight jobs above: two probabilistic, six fixed, and the decision left with the person who has to defend it.
None of that is a secret worth keeping. The reasoning is the product, and a system whose builders cannot say where their own line falls has not drawn one.
We are assembling a founding cohort now: a small number of companies willing to run a discovery analysis and let us publish the method and the results. That is the only thing that turns everything above from a position into evidence, and it is the work we would rather be judged on. If that is a trade you want to be part of, get in touch.
Common questions
Can AI be trusted to find innovation opportunities?
For generating candidates, yes. For deciding which opportunity matters most, not reliably without external controls and accountable human judgment: the failures in the published record cluster wherever the model is asked to judge rather than to produce.
Why does AI make things up?
A language model samples from a range of plausible continuations rather than retrieving a stored fact. Fluency and correctness come off the same production line, so nothing inside the output marks a sourced number as different from an invented one.
Can better prompting fix it?
The available evidence points the other way. In a field experiment with 758 consultants, the group given prompt-engineering guidance performed worse beyond the model's reliable boundary than the group given the model alone.
Does an AI innovation discovery system need to understand my industry?
Not the way a long-tenured consultant does. It needs the shape of a business of that kind: the stages where work flows and the activities supporting them. The depth about your specific company comes from you, at intake.
What makes an innovation discovery system reliable?
Roughly eight jobs, of which the model does two. The other six aim the search, carry each claim's source through to the page, grade evidence strength in code outside the model, disqualify and rank on fixed rules, and flag thin research instead of filling it in.
Sources
- Dell'Acqua, F., et al. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality." Organization Science (2026). https://doi.org/10.1287/orsc.2025.21838
- Boussioux, L., et al. "The Crowdless Future? Generative AI and Creative Problem-Solving." Organization Science 35(5) (2024). https://doi.org/10.1287/orsc.2023.18430
- Clark, J., et al. "Generative artificial intelligence use in evidence synthesis: A systematic review." Research Synthesis Methods 16–619 (2025). https://doi.org/10.1017/rsm.2025.16
- Cristofaro, M., Giardino, P.L., and Muldoon, J. "Entrepreneurial decision-making in the age of AI: Sector knowledge at the balance of intuition and analysis." Technology in Society (2026). https://doi.org/10.1016/j.techsoc.2025.103200
- Mühlroth, C., Kölbl, L., and Grottke, M. "Innovation signals: leveraging machine learning to separate noise from news." Scientometrics (2023). https://doi.org/10.1007/s11192-023-04672-y
- Ha, T., Yang, H., and Hong, S. "Automated weak signal detection and prediction using keyword network clustering and graph convolutional network." Futures (2023). https://doi.org/10.1016/j.futures.2023.103202
- Miao, H., Guo, Y., and Yuan, R. "Research on Identification of Potential Directions of Artificial Intelligence Industry From the Perspective of Weak Signal." IEEE Transactions on Engineering Management (2024). https://doi.org/10.1109/TEM.2021.3123639
- Lee, C., Kwon, O., Kim, M., and Kwon, D. "Early identification of emerging technologies: A machine learning approach using multiple patent indicators." Technological Forecasting and Social Change 127–303 (2018). https://doi.org/10.1016/j.techfore.2017.10.002
- Deloitte. 2026 Global Human Capital Trends.
- KISTI. Automated weak-signal detection report (2023). https://www.eurekalert.org/news-releases/981142
Product mechanisms described here are architecture, verified against our own repository and documentation. They are not outcome claims.