Invented specifics
Addresses, opening hours and prices arrive fully formed and completely unsourced. They read exactly like the ones that are true.
Verifiable venue intelligence
Ask an assistant where to eat and it will answer with total confidence: an address that moved two years ago, a price nobody quoted, a rating nobody gave. Power Scraping is the opposite bargain. We collect the posts you are entitled to collect, enrich them, and hand back records in which every single field cites the evidence it came from. What we cannot evidence, we return as unknown.
The problem
A language model asked to summarise a venue will produce a fluent answer whether or not the facts exist. There is no marker on the output that separates "the creator said this" from "this seemed plausible". If you are publishing a guide, briefing a client or feeding a recommendation engine, that difference is the whole job.
Addresses, opening hours and prices arrive fully formed and completely unsourced. They read exactly like the ones that are true.
A throwaway comment becomes "widely praised". A single visit becomes a rating. The chain back to who actually said it is gone by the time you see the summary.
When a client challenges one field, you have to re-do the research to answer. There is nothing to point at, which means there is nothing to defend.
The answer
The engine is the boring part. The product is the audit trail it leaves behind.
You name each profile as an authorized target and state the basis: your own account, a signed creator contract, a documented client mandate, or public-interest research — plus a reference we can audit. Nothing runs against an undeclared handle.
Collection runs through published interfaces only. Captions, frames and comments become numbered evidence items. Enrichment then proposes fields — venue, city, prices named, sentiment, experience tags — and each proposal must point at evidence ids.
Any field whose citations do not resolve is dropped, not downgraded. The release carries a provenance report (what was collected, when, under which basis) and an exception report (what was refused, skipped or discarded, and why).
A real record
This is the shape of what the API and the MCP connector return, abridged to the fields that matter here. Note the last row: the creator never gave a rating, so the field is null and says so.
rec_8d31f0a4 · from a single Instagram postmid_range| Field | Value | Evidence | Quoted from the source |
|---|---|---|---|
venue.name |
Trattoria da Oscar | Caption ev_01 |
"finally back at @trattoriadaoscar for the cotoletta" |
location.city |
Milan | Geotag ev_02 |
"Milano, Lombardia" |
price.mentions_eur |
14, 22 | Caption ev_03 |
"cotoletta €22, calice di rosso €14, worth every cent" |
assessment.sentiment |
positive | Caption ev_03 |
"worth every cent" |
venue.place_id |
ChIJ7wV0mHrBhEcR8i0s |
Resolver ev_04 |
Google Places match on name plus geotag, confidence 0.91 |
assessment.explicit_rating |
null | none | No rating appears anywhere in the post, so the field is returned as unknown rather than inferred from tone. |
An enrichment pass that cites an evidence id which does not resolve has its whole field discarded before the record is written. There is no "low confidence" tier that lets an ungrounded guess through.
Compliance posture
We run collection on your instruction, which means the limits have to be product features rather than promises. These are enforced in the API, not in a policy document.
Collection runs exclusively against profiles declared as authorized targets, each with an authorization basis and an auditable reference. An undeclared or revoked handle is refused before a single request is made.
A block, a rate limit, a login wall or a CAPTCHA ends the job. The engine does not spoof device fingerprints, solve challenges, rotate residential proxies or share session credentials — and we will not build that, on request or otherwise.
Every collection job writes an audit event naming the actor, the target and the declared basis. Acceptance of the current acceptable-use version is required before managed collection will run at all.
Collected material is yours. We act on your documented instructions, delete a target's data on request, retain for your plan's window, and never pool, sell or re-sell tenant data.
Pricing
One metered unit: a post the engine actually processed. No seat games, no per-field charges.
Free
Prove the evidence trail on a handful of posts before you commit.
Most popular
£49 /month
For one operator running continuous venue intelligence.
£399 /month
For teams collecting on behalf of named clients.
Questions
A pointer from one field of a record to one numbered evidence item collected from the source, together with the modality (caption, frame, comment, geotag, resolver), a short verbatim excerpt, the method that produced the field and its confidence. You can quote it straight back to whoever is challenging the field.
Nothing gets written. A field whose citations do not resolve against the collected evidence is discarded atomically before the record is stored. The record still arrives, with that field null, and the exception report says why it was dropped.
Ones you are entitled to have collected: your own accounts, creators you hold a signed contract with, clients who have given you a documented mandate, or subjects of public-interest research. You declare the basis and a reference per target, and you revoke it the moment it stops being true.
The job stops and reports access_blocked. We do not evade access controls
of any kind, and we will not add that capability. A refusal you can see beats data you
cannot explain.
Connect Power Scraping to Claude, ChatGPT or an IDE agent and the same evidence-grounded records become callable in a conversation: list targets, start a collection, poll it, then read records and search them. The assistant quotes the citations instead of inventing the gaps. See the quickstart.
One post the engine actually collected and enriched in the current billing period. Posts skipped by a filter, refused by an access control, or re-exported from an existing release are not metered again.
Yes. Every release is downloadable as JSON or CSV with its evidence, provenance report
and exception report intact, through the REST API or an export job.
No card. Declare one target, run one collection, and read the evidence behind every field it returns.