Original research for AI citations: what to publish and how

Original research for AI citations: which data formats ChatGPT and AI Overviews quote, SGD costs, page layout and PDPA checks for Singapore firms.

JSJun Sing Tan Updated Oct 3, 202612 min readReviewed by DMA editorial team

What you’ll learn

  • What is original research for AI citations?
  • What the evidence says, and what it doesn't
  • Why AI systems prefer to quote original data
  • Which research formats get cited
  • What each format costs in Singapore
  • How to pick a question worth answering
Original research for AI citations: a survey chart and data table being quoted inside an AI answer

Ask ChatGPT or Google's AI Mode a question about your industry and look at who it quotes. Most of the time it's a study, a benchmark or a report with a number in it. It's rarely a listicle that repeated someone else's number. If you want your brand in those answers, you need data nobody else has, published in a way a machine can lift cleanly.

What is original research for AI citations?

Original research for AI citations is data your business collects or analyses itself, such as a survey, an anonymised slice of your own records or a benchmark, published on an open page with its method, sample and date. AI answer engines cite it because they can't find the same figure anywhere else.

A model building an answer wants a source for each claim. When ten sites repeat one statistic, the strongest domain usually gets the credit. When your page is the origin of the number, the model has to come to you, and so does every writer who quotes it later.

2.7%of AI-cited pages were primary research (Growth Memo sample)
8%+of all citations went to those few pages
40%max visibility gain from GEO edits in the Aggarwal et al. study
115.1%visibility lift for rank-5 sources that cited their sources

What the evidence says, and what it doesn't

The GEO paper

The first is GEO: Generative Engine Optimization by Pranjal Aggarwal and colleagues at Princeton, IIT Delhi and the Allen Institute, presented at KDD 2024. They built a benchmark of 10,000 queries, rewrote source pages in different ways and measured how much of each page showed up in generated answers. Three edits beat the rest: adding statistics, adding quotations, and adding citations to sources, in roughly that order of usefulness for most queries. Their abstract puts it plainly:

"Through rigorous evaluation, we demonstrate that GEO can boost visibility by up to 40% in generative engine responses."

Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024

Keyword stuffing, the old SEO reflex, did worse than leaving the page alone on their main metric. The more useful finding for a smaller brand sits in their ranking table. Pages that ranked lower in the search results gained the most from these edits, while the top-ranked source often lost share.

Edit tested in the GEO paperChange for the rank-1 sourceChange for the rank-5 source
Cite sources-30.3%+115.1%
Add quotations-22.9%+99.7%
Add statistics-20.6%+97.9%
Relative change in visibility by search ranking position. Source: Aggarwal et al., Table 2.

Treat this as a lab result. The engines they tested in 2023 have since been replaced, and the authors found the effect changed from one subject area to the next. What we see on client pages in 2026 points the same way, though: a specific number with a source beside it gets pulled into answers far more often than a general claim.

The Growth Memo citation study

The second is a July 2026 piece in Growth Memo where Kevin Indig and Amanda Johnson worked through citation data from Gauge, and Arcalea's summary sets out the figures if you don't have a subscription. Their sample covered 301 cited pages across 316 prompts in seven industries, which together drew 1,075 citations. Of those pages, only 8 were genuine primary research, roughly 2.7% of the set, yet they pulled in 90 citations between them. That's just over 8% of the total and 3.3 times the citation density of the other pages.

Here's the catch most summaries skip. 75 of those 90 citations went to one cluster: benchmarks that ranked named vendors on speed and cost. Research written up as a narrative report barely registered. So original data isn't enough on its own. The format decides whether a model can use it.

What this means for you

A benchmark that answers "which option is best for X" beats a 40-page PDF with a nicer cover. If your study can't be turned into a ranked comparison or a single quotable figure, rethink the question before you spend on fieldwork.

Need help with marketing? DMA builds and runs campaigns that grow Singapore businesses.

Free strategy call ›

Why AI systems prefer to quote original data

A few plain reasons push a model toward a primary source.

First, uniqueness. Every blog that later quotes your number links back to you, and the model sees those links too.

Second, verifiability. A page that shows how many people were asked, when, and how, gives the system and the reader something to check. A number with no method beside it looks like marketing, because it usually is.

Third, freshness. Questions like "how much does X cost in 2026" need a current answer. The newest credible study in a category tends to replace older ones, which is why a study you repeat every year compounds. Last year's edition keeps its links, this year's edition takes the citations.

Google's own guidance is less exciting than the vendor pitch. Its page on AI features says there are no extra requirements to appear in AI Overviews or AI Mode, and no special AI text files or markup needed. It asks for the basics, such as keeping important content in text form and making sure structured data matches what's visible on the page. So the work on a research page is mostly about clarity and trust, which you control.

Which research formats get cited

You don't need a six-figure annual report. Most SMEs already sit on data worth publishing. These are the five formats we use, from cheapest to most involved.

Anonymised first-party data

Your booking system, CRM, ecommerce store or helpdesk already records patterns nobody else can see. A clinic knows which weekday has the most no-shows. A renovation firm knows the median quote for a 4-room HDB kitchen across 300 jobs. Aggregate it, strip anything personal, and you have a statistic that's yours by default.

Customer or industry survey

Surveys capture opinions and self-reported behaviour that your records can't. Run one to your own list for low cost, or pay a panel provider when you need a broad sample of Singapore consumers. Write the question list and the method section before you send anything. That discipline stops you from asking leading questions you'll regret when a journalist reads them.

Public data analysis

This is the most underused option here. SingStat Table Builder and data.gov.sg publish thousands of datasets that few marketers ever touch. Cross two of them, or slice one by a cut nobody has published, and the analysis is original even though the raw data is public. Cite the government source on every chart.

Benchmark index

Test a fixed set of products, providers or sites the same way, then rank them. This is the format the Growth Memo study found doing most of the work. It answers comparison questions directly, which is exactly what people type into ChatGPT. Publish the scoring rules so competitors can't dismiss the result.

Expert panel

Put the same questions to 10 to 20 named practitioners and report where they agree. You get fewer hard numbers, so these pages win "how should I" questions more than "how many" ones, and every panellist has a reason to share the result.

What each format costs in Singapore

None of the guides we read gave real costs, so here are ours. These are planning estimates from scoping this work for local SMEs, not quotes, and they assume you write and design in-house or with your agency. Panel prices in particular move with sample size and how niche your audience is.

FormatEstimated cost (SGD)Time to publishBest for
Anonymised first-party dataS$1,500 to S$6,000 (analyst and writing time)2 to 6 weeks"How much" and "how often" questions in your niche
Survey to your own listS$500 to S$3,000 (tool, incentives, analysis)4 to 8 weeksB2B opinions, customer behaviour
Paid panel survey, Singapore consumersS$8,000 to S$25,0006 to 10 weeksPress pickup, national claims
Public data analysisS$1,000 to S$5,0002 to 4 weeksTrend pieces, local market sizing
Benchmark indexS$3,000 to S$15,000 first edition, less to refresh4 to 8 weeks"Which is best" comparison prompts
Expert panelS$2,000 to S$8,0004 to 8 weeksStrategy questions, relationship building
DMA planning estimates as of October 2026, to be confirmed against your own sample size, incentives and design work.

Our usual advice for a first study: start with first-party data or public data analysis. Both are cheap and quick, and they show you what your audience reacts to before you commit five figures to a panel.

How to pick a question worth answering

The best research question is a missing statistic. Look for a number your industry repeats that traces back to nothing, an old overseas study, or a vendor blog quoting another vendor blog. Replace it with a current Singapore figure and you become the citation.

Read your own top pages and mark every claim that starts with "most", "many" or "on average" without a source. Then put your customers' questions to ChatGPT and Google AI Mode and note where the answer is vague or out of date. Questions starting with "how much", "how many" and "which is best" are the strongest candidates, because the answer is a number or a ranking.

Then check you can answer it honestly. If your sample can't support the headline you want, rewrite the headline and leave the sample alone.

How to structure the data page so AI can extract it

This is where most studies lose. The data is good, the page buries it. A model reads in chunks, and each chunk should make sense on its own. Kevin Indig's February 2026 analysis of 1.2 million search results described AI as reading like a busy editor rather than a patient student, so the page has to work for someone who skims the top and leaves.

Lead with findings, one stat per sentence

Put your three to five headline findings in the first screen. Each finding gets its own sentence, and each sentence carries the number, the unit, the sample and the date. If a reader copies that one sentence into an email, it should still be accurate.

Hard for AI to quoteEasy for AI to quote
Most of our customers prefer weekend appointments.62% of 1,140 bookings at our 3 Singapore clinics between January and June 2026 were made for Saturday or Sunday.
Costs have gone up a lot recently.The median kitchen renovation quote in our 2026 sample of 312 HDB jobs was S$14,800, up from S$12,900 in 2025.
See chart below for results.A HTML table listing each finding, its sample size and its collection dates.
Made-up example sentences, written for this guide to show the pattern.

Show the method on the same page

Add a methodology section with the sample size, who was included, how they were recruited, collection dates, the questions asked and known limits. Put it on the same page as the findings. When the method sits in a PDF behind a download form, nobody checks it, and that includes the crawlers.

Use HTML tables, not chart images

A bar chart exported as a PNG is invisible to most text extraction, so publish the numbers behind every chart as an HTML table directly below it, with clear headers and units. People still like the chart, and the table is what gets quoted. Google's advice to keep important content in text form matters more on a data page than on any other page you own.

Date everything and update in place

Show a published date and a last-updated date near the top. When you refresh the study, update the same URL and add a short change log, so the links you earned last year keep working. Put the year in the title only if you plan to update it every year.

Mark it up honestly

If you publish the raw dataset, Google's Dataset structured data describes it for Dataset Search. Article and FAQ markup help on the write-up. Whatever you mark up must match the visible text. Markup doesn't make a weak study citable, and Google says none is required for AI features.

Gated or open?

Keep the findings, tables and method open. Gate the extras if you want leads, such as an industry cut, a raw CSV or a slide deck. A model can't cite what it can't read, and neither can a journalist.

PDPA checks before you publish customer data

None of the three guides we compared mention this. It's also the one part of a research project that can land you in front of a regulator, because a study built from Singapore customer records stays under the Personal Data Protection Act until the data is properly anonymised.

PDPC's Guide to Basic Anonymisation sets out five steps for simple datasets. You get to know the data, strip direct identifiers, apply anonymisation techniques, work out the risk that someone could be re-identified, and then manage whatever risk is left. Of the techniques it describes, aggregation is the one a marketing study leans on most, since you'll be publishing medians and percentages rather than individual rows.

Small groups are where studies leak. A line like "Our only client in Tuas with more than 200 staff cut ad spend by half" identifies a company, and maybe a person. Merge or drop any segment small enough that a reader could work out who it is. Check that your original collection notice covers research use. With surveys it's simpler: say on the first screen that aggregated results will be published.

The stakes are real. PDPC's enforcement update explains that since 1 October 2022 the Commission can fine an organisation with Singapore turnover above S$10 million up to 10% of that turnover, and others up to S$1 million. No marketing study is worth that exposure, so if you're unsure about any segment, get your data protection officer to sign off before launch.

Getting the data cited: distribution and link building

A study nobody sees earns nothing. AI systems lean on sources that other sources cite, so your distribution plan is also your citation plan, and data studies remain one of the few reliable ways to earn editorial links.

Pitch two or three findings to the journalists and newsletter writers who cover your sector, each with a ready-to-paste sentence and a link to the method. Then go back to your own site. Add the new statistic to your service pages, product pages and older blog posts that make related claims, and link each one to the study. Original data works hardest on the pages that sell, because a sourced number persuades better than an adjective.

If you want the study to support a wider link building push, keep one evergreen URL for the research and send every pitch there. Splitting links across a PDF, a press release and a blog post wastes them.

How to measure whether it worked

Write down 20 to 30 prompts your study should answer before you publish. Run them in ChatGPT, Perplexity, Gemini and Google AI Mode and note whether each answer cites you, links to you or only paraphrases your figure. Do this once a month with the same prompt set, otherwise the trend tells you nothing.

Also watch referring domains to the study URL and AI referral traffic in GA4. In our experience the links and mentions arrive first and the AI citations follow some weeks later, once enough other pages point back to you. If nothing moves after a quarter, check the format before you blame the data.

For the wider picture of how Google's AI answers behave, our roundup of Google AI Overviews statistics tracks the numbers, and our explainer on generative engine optimisation covers the rest of the playbook. Research works best inside a planned content marketing strategy, where each study feeds a quarter of posts, pages and pitches.

Frequently asked questions

Can original data really get my brand cited by ChatGPT and AI Overviews?

It improves your odds, with no guarantee. The GEO paper found that adding statistics and sources raised visibility in generated answers by up to 40%, and the Growth Memo sample showed primary research earning far more citations per page than ordinary content. Format and distribution still decide the outcome.

How big does my survey sample need to be?

Big enough to support the claim you make. A niche B2B study can be credible with a few hundred qualified respondents if you state the sample and who they are. A national consumer claim needs a properly sampled panel. Always publish the sample size beside each finding.

Do I need schema markup for AI to cite my research?

No. Google says no special markup or AI text files are needed for AI Overviews or AI Mode. Dataset markup helps your raw data appear in Google Dataset Search, and any markup you add must match the visible page. Clear text and HTML tables matter more.

Can I publish data from my Singapore customers?

Yes, once it's properly anonymised and aggregated. Follow PDPC's Guide to Basic Anonymisation, remove direct identifiers, merge small segments that could identify someone, and check that your collection notice covers research use. Ask your data protection officer to sign off.

Should I gate my research behind a form?

Keep the findings, tables and methodology open so AI systems and journalists can read and cite them. Gate optional extras such as raw files, industry breakdowns or slide decks if you want leads from the study.

How often should I repeat a study?

Once a year suits most benchmarks and surveys. Update the same URL, keep earlier editions linked from it, and add a change log. A repeated study builds a time series that newer sources cannot copy.

Want a study that AI and journalists will quote?

We help Singapore businesses find the missing statistic in their market and turn it into a study worth quoting. Our AI SEO team runs the data page and the monthly citation tracking. Our content marketing agency team takes the same study and spreads it across three months of posts, pitches and page updates. Talk to us about your first study.

JS

Jun Sing Tan

Jun Sing Tan is part of the content team at D’Marketing Agency, a Singapore digital marketing agency specialising in SEO, SEM, social media & lead generation. About DMA ›

Get weekly SEO & marketing tips

Join 3,000+ marketers growing with D’Marketing Agency.

Share f X in

Want results like this for your business?

D’Marketing Agency helps Singapore brands grow with SEO, Google Ads, social & lead generation.

Get a free strategy call
Book your discovery call

Important Notice: Protect Yourself from Scammers

We have been informed that scammers are impersonating D’Marketing Agency Pte. Ltd. Please note that our official website is / , and phone number is +65 8971 9405

Always verify contacts before engaging and do not share personal information with unverified sources. If you have been scammed or approached by scammers impersonating us, please report the number to the platform you are using. We are not responsible for these fraudulent activities.

For any concerns or to confirm the legitimacy of any communication, contact us directly through our official phone number or website. Your safety and security are our top priorities.

Thank you and stay safe!