How to Do Serious Research in the AI Era
Digital research has always been a blurred practice. The word gets attached to work that has not earned it.
A survey collecting answers through a web page is not digital research. It is a survey with a web form in front of it. The collection method went digital. The research did not.
AI added a new version of the same confusion. A model that crawls hundreds of sources and returns a written brief gets called deep research. It is deep search. It reads what is already published and reports back what it found. And most of what sits on the web is second-party or third-party data: someone else’s collection, gathered for someone else’s question, shaped by someone else’s design. Reading a lot of it quickly is retrieval. It is not research.
Collating a series of those secondary sources into a summary is editorial compilation. It stays compilation until someone does the surgical part: pulling the contextual signals out of the raw data and turning them into measured differences in meaning. Skip that step and you have gathered sources, not studied anything.
That surgical step is the subject of this article. Not the tool. It is the method that researchers need to learn.
Correlation is easy. Dimension is the work.
For years I trained people in web analytics, and one principle sat at the front of every session: do not measure in straight lines.
A single timeline is a straight line. It shows one thing over time, and all you can read is the difference between one point and the next. Traffic rose in March, fell in June. That is linear measurement, and it stays shallow, because a line can only describe itself.
Add one variable and everything opens up. Put a second timeline beside the first. Sales against sentiment. Sign-ups against price. Now you can read three things instead of one: how the first line moves, how the second line moves, and how the two differ at the same point in time. One variable turned a line into a dimension. That two-line read was the first thing I taught, because it is where measurement stops being flat and starts having depth.
When you work from a compiled summary, you rarely get past the single line. You get “these two things seem related,” which is correlation, the easiest observation to make. What you cannot do is measure the relationship, because the raw data that would let you model it was never handed to you. Noticing two lines move together is not the same as measuring how, how much, and when the pattern breaks.
What it gets right
None of this makes compilation worthless.
A skilled editor can take those curated sources and produce something genuinely useful. Direction. A first map of a market. A briefing that saves a team a week of reading. With judgment applied on top, secondary research can be insightful and can occasionally change a decision.
The problem is not that it is useless. The problem is its limitation.
The limitation
Secondary research carries whatever bias came with it. The question someone else asked. The sample they chose. The angle they were paid to take. You absorb all of it the moment you cite the summary, and you usually cannot see it, because the assumptions are buried underneath.
Depth suffers too. A summary compresses. It keeps the headline and drops the exceptions, and the exceptions are often where the real insight lives.
There is a quieter cost. Trusting a cited answer because it is cited is the same reflex people once had for a government index or a famous study. Authority is not the same as validity. A source can be well known and still be the wrong source for your question.
Where serious research begins
Serious research starts one step earlier than the compiled summary lets you start.
It starts with data you gathered yourself, for your own question, before anyone summarized it.
The strongest version of this is behavioral data. What people actually did when no one was watching and no one was asking. A search query is behavioral. So is a click, a purchase, a moment someone gave up and left. Nobody framed it for a study. The person framed it themselves, for themselves. That honesty is something no commissioned report can give you.
When you own that data, AI stops being an answer machine and becomes something more valuable. A researcher’s lens.
What building an AI research platform taught me
I built an AI platform to do this at a depth no spreadsheet reaches.
Plain text search measures how often a term appears and ranks by that count. It is shallow. The platform goes past word counting to nearest-neighbor discovery, finding the records that sit closest in meaning even when they share no words. Each method like this adds another dimension to read across, and reading across dimensions is where discovery happens.
A table cannot hold that for long. A table is a two-dimensional matrix: rows, columns, a flat grid where relationships are read one pair at a time. So I do not stop at the table. I model the data as a graph. Every record becomes a node. Every relationship becomes a link between nodes. A node connects to many others, and those others connect onward, until the data stops being a grid and becomes a network you can move through. A flat table tells you that A relates to B. A graph lets you follow A to B to C, and reach a connection no pairing of columns would surface.
This is the two-line principle taken to its end. One line is flat. Two lines make a dimension. A graph makes a space: a data universe of connected nodes, where a relationship can run in any direction, and depth is measured by how far the connections reach rather than how many columns you can fit.
Depth is worth nothing without accuracy, so nothing is taken on trust. Every pattern the model reads is checked against the actual records, matching its meaning-based reading against exact counts in a structured database, confirming the data is really there and not inferred. One method finds what the data means. The other proves it exists. A finding only stands when both agree.
No human working data by hand reaches this depth at this scale. That is not a criticism of the analyst. It is arithmetic. A person can hold a few relationships in mind at once. The machine can weigh hundreds and check every one against the source.
The methodology is the point
The technology is not the center of this. The methodology is.
The algorithms create the dimensions. The researcher decides which ones matter, what question they answer, and whether a pattern means anything at all. AI did not replace that judgment. It extended the reach of it.
This is what augmentation truly is. Not a faster tool. A wider analytical mind. The human still does the thinking. The machine lets that thinking travel further than one person could carry it alone. Call it a technical advance and you have missed what changed. What changed is how far a serious researcher can now see.
You can borrow the finding, not the practice
AI did not lower the standard for research. It lowered the effort, which is not the same thing. Effort dropping is only good news if the standard holds.
Secondary research can borrow a serious researcher’s findings. It cannot borrow the practice that produced them. The finding is the part you can quote. The practice is everything underneath it: the question framed before the tool was opened, the data gathered first-hand, the dimensions tested against each other, the results checked against the source until they held. None of that travels with the citation. You get the answer without the discipline that earned it.
That is the difference between borrowing a conclusion and standing behind one.
I am serious about research. I choose the serious side, and I always have. The tools will keep changing. My choice will not.


