Hacker News Scraper
Search Hacker News stories, Show HN, Ask HN, comments, or the front page by keyword and get clean JSON with points, author and comment links.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
query,tags,sortBy(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.001 per item = $1 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Item returned | Charged per HN story/comment returned. | $0.001 |
Pay-per-event pricing: you are billed per result, not per subscription — a run that returns nothing costs nothing beyond the start fee. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-12, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
query | Keywords to search Hacker News for (e.g. "openai", "rust async"). Leave empty to get the newest/front-page items for the chosen tag. | string |
tags | Which kind of Hacker News items to return: stories, Show HN, Ask HN, comments, or current front page. | string |
sortBy | Order results by Algolia relevance (best match) or by date (newest first). | string |
minPoints | Only return items with at least this many points (upvotes). Leave empty for no minimum. Applies mainly to stories. | integer |
maxItems | Maximum number of items to return. The actor paginates the API until it has this many (or runs out). | integer |
notionConnector | Optional. Write each item as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default) — results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
authorcreatedAthnUrlnumCommentsobjectIdpointstexttitletypeurlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
3 ready-to-run use cases
Search Hacker News Stories by Keyword: OpenAI & More
Search Hacker News for any keyword like OpenAI and get matching stories ranked by relevance, each with points, author, and a direct comment-thread link.
Scrape New Show HN Launches: Newest Products First
Founders and marketers tracking competitors get the latest Show HN posts, newest first, with product titles, points, and links to every fresh launch.
Hacker News Front Page Scraper: Live Top Stories
A snapshot of the current Hacker News front page: titles, points, authors, and comment counts for every story ranking right now, returned as structured JSON.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Research MCP Server — 10 Tools for AI Agents
One MCP server URL gives Claude, Cursor or ChatGPT 10 research tools: arXiv, Reddit, GitHub, HN, OpenAlex, Wikipedia, CoinGecko, GDELT and more.
DEV.to Scraper
Scrape DEV.to articles by tag, author or sort: title, URL, tags, reactions, comments, reading time, cover image and full body.
Wikidata Scraper
Resolve names or Q-ids to Wikidata entities for $0.20 per 1,000 plus $0.001 per run. Labels, aliases, instance-of, claims and Wikipedia links.
Domain Inspector
Bulk-inspect domains: DNS records, RDAP registrar and expiry, redirects, TLS dates, security headers, robots, sitemaps and tech detection.
GitHub Scraper
Search GitHub repos or users and export clean rows: stars, forks, language, topics, license, plus user bio, company, location and follower count.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tags. Get structured questions with score, answers, views, tags, author, and link as JSON or CSV.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- Hacker News
Hacker News Scraper
Scrape Hacker News via the public Algolia HN Search API — no API key, no login, no anti-bot. Search stories, Show HN, Ask HN, comments, or the current front page; sort by relevance or date; filter by minimum points.
What it does
- Queries
https://hn.algolia.com/api/v1/search(relevance) orsearch_by_date(newest first). - Paginates automatically until it has collected
maxItems. - Returns clean, normalized JSON for each item (HTML stripped from story/comment text).
- Deduplicates by Hacker News
objectID.
Input
| Field | Type | Default | Description |
|---|---|---|---|
query | string | "openai" | Keywords to search for. Empty = newest/front-page items for the chosen tag. |
tags | select | story | story, show_hn, ask_hn, comment, or front_page. |
sortBy | select | relevance | relevance or date (newest first). |
minPoints | integer | — | Only items with at least this many points. |
maxItems | integer | 50 | Max items to return (1–1000). |
proxyConfiguration | proxy | off | Optional; the public API has no anti-bot, so no proxy is needed. Only enable it if you hit IP rate limits at very high volume. |
> front_page note: the front_page tag always returns the ~30 items currently on the HN front page (no keyword search). For topic searches use story/show_hn/ask_hn/comment instead.
Output
Each successful item:
{
"ok": true,
"objectId": "44159823",
"type": "story",
"title": "OpenAI ...",
"url": "https://example.com/article",
"author": "someuser",
"points": 412,
"numComments": 188,
"createdAt": "2026-06-01T12:34:56.000Z",
"text": "",
"hnUrl": "https://news.ycombinator.com/item?id=44159823"
}
For comments, type is comment, title/url reflect the parent story, and text is the (HTML-stripped) comment body.
Some fields can be null depending on the item: url (Ask HN/text posts and self-posts have no external link), points and numComments (often absent on comments), and text (empty string for link stories). title, author, objectId and hnUrl are present for well-formed items but defensively default to null if the API omits them.
On failure or no results, a single diagnostic row is emitted with ok:false and an errorCode (e.g. NO_RESULTS, RATE_LIMITED, NETWORK) — and nothing is charged.
Delivery
Results always land in the run's dataset (JSON, CSV, Excel, XML via the API or the Console). Two optional inputs, notionConnector and notionParentId, will additionally write one Notion page per item when the run finishes — authorize a Notion MCP connector once under Settings → API & Integrations → MCP connectors, then pick it in the input form. Leave them empty to skip.
proxyConfiguration is off by default. The Algolia HN Search API is public and has no anti-bot, so a proxy adds latency and cost for nothing; only switch it on if you are running enough volume to hit IP-level rate limits.
What people use it for
- Brand and competitor monitoring — search
storyandcommentfor your product name and watchpoints/numCommentsto see which threads are actually getting traction. - Launch tracking —
show_hnsorted bydategives you every new Show HN as it lands. - Hiring and demand signals —
ask_hnsurfaces "Who is hiring" and "Who wants to be hired" threads. - Front-page snapshots —
front_pagereturns the ~30 items on the HN front page right now, for a daily digest or a dashboard. - Filtering the noise —
minPointsdrops everything below a score threshold before you are billed for it.
Pricing
$1.00 per 1,000 items ($0.001 each), with no run-start fee. Flat rate — no volume tiers, no plan gates — and you are charged only for items actually returned.
One charge per returned item (item event). Diagnostic and empty rows are never charged, so a query that matches nothing costs nothing. Because you pay per row, set maxItems to what you actually need — the default is 50 and the ceiling is 1,000 per run.