Fetch and Extract an X Post

Site x.comTask fetch-x-postVersion v13Updated Sep 20, 2026Category social

Load an X status URL, or resolve a relevant post from an account timeline, then deterministically extract timestamp, post text, visible engagement counts, attached image and video metadata, page signals, and visible access restrictions. This skill was captured from a live agent session on x.com and is published here as a reusable recipe for agents.

NoteSelectors and URL schemes drift as sites change. A skill is a snapshot of what worked when it was captured, not a contract — agents re-learn it when it stops working.

Fetch an individual X post and return whether it was accessible, its ISO and displayed timestamp, text, visible engagement counts, attached image and video metadata, page signals, and visible restriction or error signals. When only an account and a content clue are available, resolve candidate status URLs from the account's /media or /with_replies route before opening the selected post. Preserve enough structured information to avoid confusing replies or related posts with the requested post.

Use Cases

Use when the caller supplies an X status URL or opaque status ID. It also applies when a relevant post must be found among an account's media posts or surrounding posts using text, timestamp, attached-image, or neighboring-post clues. This recipe is for reading a single post and its directly attached media and visible metrics; it is not for broad timeline searches or collecting replies.

Automation Flow

  1. If a status URL or status ID is available, navigate directly to https://x.com/{username}/status/{status-id}?s=20 (or use the caller-provided status URL), waiting for domcontentloaded. Do not visit the homepage or use the search box first.
  2. If no status ID is available, navigate directly to https://x.com/{username}/media for image/media clues, or https://x.com/{username}/with_replies when surrounding posts or neighboring context is needed. Run the candidate evaluator below, select the candidate whose text, timestamp, image alt text, or neighboring context matches the clue, and read its opaque /status/{status-id} URL from statusUrl; never guess the ID. Navigate directly to that status URL and run the post evaluator.
  3. On the loaded status page, run the post evaluator below in the same browser call. Use the first article as the requested post; later articles are replies or related content. The time element supplies the timestamp used to verify publication time.
  4. Combine the evaluator result with browser navigation status. Report tool-level failures separately from page-level restrictions; successful navigation without a matching article is not successful extraction.

Post evaluator:

(() => {
  const bodyText = (document.body?.innerText || '').trim();
  const articles = [...document.querySelectorAll('article')];
  const metricPattern = /reply|repl(?:y|ies)|retweet|repost|like|bookmark|view/i;
  const posts = articles.map((article, index) => {
    const tweetText = [...article.querySelectorAll('[data-testid="tweetText"]')]
      .map(n => (n.innerText || n.textContent || '').trim()).filter(Boolean);
    const time = article.querySelector('time');
    const links = [...article.querySelectorAll('a')]
      .map(a => ({text: (a.innerText || '').trim(), href: a.href || ''}))
      .filter(x => x.text || x.href.includes('t.co')).slice(0, 30);
    const imgs = [...article.querySelectorAll('img')]
      .map(img => ({alt: img.alt || '', src: img.currentSrc || img.src || '', width: img.naturalWidth || null, height: img.naturalHeight || null}))
      .filter(x => x.src).slice(0, 20);
    const vids = [...article.querySelectorAll('video')]
      .map(video => ({src: video.currentSrc || video.src || '', poster: video.poster || '', width: video.videoWidth || null, height: video.videoHeight || null, duration: Number.isFinite(video.duration) ? video.duration : null, currentTime: video.currentTime || 0, paused: video.paused}));
    const engagement = [...article.querySelectorAll('button, [role="button"], [data-testid="reply"], [data-testid="retweet"], [data-testid="like"], [data-testid="unretweet"], [data-testid="unlike"], [data-testid="bookmark"], [data-testid="app-text-transition-container"]')]
      .filter(node => !node.closest('[data-testid="tweetText"]'))
      .map(node => ({testid: node.getAttribute('data-testid') || null, label: node.getAttribute('aria-label') || node.getAttribute('title') || '', text: (node.innerText || node.textContent || '').trim()}))
      .filter(x => metricPattern.test([x.testid, x.label, x.text].filter(Boolean).join(' ')))
      .filter((x, i, all) => x.label || x.text || x.testid && all.findIndex(y => y.testid === x.testid && y.label === x.label && y.text === x.text) === i)
      .slice(0, 30);
    return {index, text: (article.innerText || '').trim(), tweetText, timestamp: time ? {datetime: time.getAttribute('datetime') || null, text: (time.innerText || '').trim()} : null, links, imgs, vids, engagement};
  });
  const chromeText = (() => {const clone = document.body?.cloneNode(true); if (!clone) return ''; clone.querySelectorAll('article').forEach(node => node.remove()); return (clone.innerText || '').trim()})();
  const restrictionPatterns = [/log in to x/i, /sign up for x/i, /create an account/i, /something went wrong/i, /this page doesn.?t exist/i, /post isn.?t available/i, /content warning/i, /account suspended/i, /rate limit exceeded/i, /unusual activity/i, /complete the challenge/i, /are you a robot/i];
  const restrictions = restrictionPatterns.filter(p => p.test(chromeText)).map(p => p.source);
  const primary = posts[0] || null;
  const hasText = Boolean(primary?.tweetText?.length);
  const hasMedia = Boolean(primary && (primary.imgs.length || primary.vids.length));
  return {url: location.href, success: Boolean(primary) && (hasText || hasMedia), hasText, hasMedia, extractedPostText: primary?.tweetText?.join('\n') || null, timestamp: primary?.timestamp || null, engagement: primary?.engagement || [], posts, media: primary ? {images: primary.imgs, videos: primary.vids} : {images: [], videos: []}, restrictions, pageTitle: document.title || null};
})()

For resolving candidates from /media or /with_replies, use this evaluator on that page:

(() => [...document.querySelectorAll('article')].map((article, index) => {
  const statusUrl = [...article.querySelectorAll('a[href*="/status/"]')].map(a => a.href).find(href => /\/status\/\d+/.test(href)) || null;
  const text = (article.innerText || '').trim();
  const tweetText = [...article.querySelectorAll('[data-testid="tweetText"]')].map(n => (n.innerText || n.textContent || '').trim()).filter(Boolean);
  const time = article.querySelector('time');
  const images = [...article.querySelectorAll('img')].map(img => ({alt: img.alt || '', src: img.currentSrc || img.src || '', width: img.naturalWidth || null, height: img.naturalHeight || null})).filter(x => x.src);
  const videos = [...article.querySelectorAll('video')].map(video => ({src: video.currentSrc || video.src || '', poster: video.poster || '', width: video.videoWidth || null, height: video.videoHeight || null, duration: Number.isFinite(video.duration) ? video.duration : null}));
  return {index, statusUrl, text, tweetText, timestamp: time ? {datetime: time.getAttribute('datetime') || null, text: (time.innerText || '').trim()} : null, images, videos};
}))()

If media metadata is incomplete, optionally open a[aria-label="View media"], wait briefly, and rerun the post evaluator. Attached photo routes may also be inspected directly at https://x.com/{username}/status/{status-id}/photo/{n}. For a video whose meaning is not represented by text or metadata, inspect a small set of temporal positions such as 0%, 25%, 50%, 75%, and near the end; do not infer visual claims from duration or poster metadata alone.

Possible Friction Points

  • The primary post is represented by an article; its text is normally in [data-testid="tweetText"]. Additional articles may be replies or related content, so use the first article for a requested status page.
  • The primary article's time element exposes an ISO datetime and often a displayed relative timestamp. Preserve both; use ISO datetime for timestamp verification rather than relying on relative text.
  • Visible engagement controls commonly expose counts in aria-label, title, button text, or data-testid; preserve the raw engagement entries rather than assuming a fixed control order or parsing localized abbreviations.
  • Account media posts are reachable at https://x.com/{username}/media, while surrounding posts and replies are reachable at https://x.com/{username}/with_replies. Status links embedded in article cards expose the opaque ID needed for a direct status URL; resolve and read that link rather than deriving or guessing the ID.
  • Images attached to a post are exposed as article img elements. Preserve their alt, src, and natural dimensions; photo-specific routes use /photo/{n} when direct inspection is necessary.
  • Videos may not expose a usable video element until playback or the media viewer is opened. Preserve currentSrc/src, poster, intrinsic dimensions, duration, and playback state. For detailed explanations, sample a few meaningful timestamps rather than repeatedly seeking through the entire video.
  • X may render a login wall, unavailable-post message, consent screen, rate-limit page, or anti-bot challenge without exposing tweet text. Return success: false and the visible signals in that case.
  • The status ID is opaque and must come from the caller's URL or be resolved from a status link in an account timeline; never guess it from the username.
  • The s=20 query parameter is optional for direct status navigation and should be preserved when supplied.
  • A browser navigation status is a tool-level result, while success indicates that the primary post was actually extracted from the DOM. An X post can carry media with no text at all, so success is true when either text or media was extracted, and hasText and hasMedia report which. Requiring text alone would report an accessible image or video post as a failed extraction.
  • Restriction patterns are tested against the page text outside the article elements, and they match full wall phrasing. A public post's own words otherwise trip them: /log in/ and /sign up/ match the logged-out banner on every post, and /challenge/, /robot/, and /rate limit/ match ordinary post text, leaving restrictions non-empty on nearly every page.
  • Engagement candidates are restricted to controls and the known metric test IDs, and nodes inside [data-testid="tweetText"] are excluded. Scanning every [data-testid] sweeps the post's own words into the metrics: a post reading "I like this" matches the metric pattern and becomes an engagement entry.

Call it

GET https://production-sfo.browserless.io/skills?token=TOKEN-HERE&domain=x.com&task=fetch-x-post