Pattern-of-life analysis, cross-platform correlation, social network analysis, and mis/dis/malinformation.
Social Media Intelligence (SOCMINT) is the OSINT discipline that covers what people, groups, and organizations publish on social platforms. Where the previous chapters focused on search engines and one-off identification, SOCMINT turns a person's continuous public posting into structured knowledge: their identity, home base and workplace, daily rhythm, associations, and degree of exposure.
Passive vs. active collection: passive (observing public content with no interaction) is the default and the safe posture. Active (following, messaging, joining a group) is intrusive, frequently outside scope, and exposes the investigator to OPSEC risk. Legally and ethically: use only publicly accessible material, never circumvent access controls, and remember that SOCMINT on real people is personal-data processing under GDPR, and it needs a lawful basis and proportionality.
Mundane posts leak more reliably than deliberately personal ones. A post that's obviously about someone's private life gets written with at least some awareness that it's revealing, since people self-censor, however imperfectly. A completely ordinary post (asking for a lunch recommendation near an office, complaining about a slow commute, mentioning being at a desk on a weekday afternoon) carries no such filter. The poster isn't thinking about OPSEC when writing something that mundane. One throwaway lunch-recommendation post can quietly answer a work-area location, an employment status, and a rough daily schedule all at once, precisely because it never felt like the kind of post worth being careful about.
Instead of asking "what did this person say?", ask "what does the pattern of what they say reveal?" From social media alone you can often reconstruct: a daily rhythm (wake/commute/work/home, inferred from posting times and content), a weekly rhythm (which days break the pattern: a regular free afternoon, an evening class), movement (home district vs. work district vs. travel, from location clues over time), and life events (a new job, a move, a graduation). The power comes from repetition. One 8:05am post means nothing — three commute-timed posts across different weeks establish a reliable morning window.
Timestamps are one of the richest, most under-used SOCMINT signals. Platforms often display localized time while keeping an absolute UTC timestamp underneath: if the post shows 6pm but the raw timestamp reads 3:02pm UTC, that gap is itself a +3 timezone tell. Never rely on one indicator alone: combine the shape of the active-hours histogram, any locale/timezone fields embedded in the page source, and content references to local events. Caveat: post scheduling, cross-posting tools, VPNs, and travel can each distort timestamp analysis, so treat timestamps as supporting evidence, never as proof.
In practice, start by mapping every available timestamp onto a 24-hour axis, split by day of the week. The activity clusters (morning, midday, evening) and the long dead zones (local night) sketch the template of the subject's day before you've read a single word of content. Never call a timezone off one post: read the shape of the distribution across dozens of posts, and want convergence from at least two independent categories of evidence (behavior plus metadata, or behavior plus content). When you do report a timezone finding, explicitly note the factors that could invalidate it, such as possible post scheduling, a VPN, or recent travel.
Location leaks through several channels: explicit geotags/check-ins, landmarks and place names in text, photos (see Chapter 6), commute references (a station or line name narrows home/work), and repetition over time. Geotags and check-ins are usually the single strongest clue, since people forget these are attached. Repetition also matters: a home area shows up week after week, while a place visited once reads as travel. The "read then verify" discipline matters here: a named hill is a lead, not an answer, until a map lookup confirms it. Verification follows a fixed chain: toponym → map/gazetteer lookup → neighborhood or transit corridor → compatibility check against the subject's other geographic traces. A station name isn't a finding until it's resolved to a line, and the line to the neighborhoods it serves. A new element only enters the findings once it's consistent with the picture you already hold, or convincingly explains the deviation.
Real people are active on more than one platform, and the fullest picture emerges when you merge the fragments from each. The most common pivot is handle reuse: starting from one known handle, you enumerate other platforms for it, then confirm it's genuinely the same person, not a coincidental match. Corroborating signals include the same avatar or color scheme, a consistent join date, a consistent bio tone, and content referencing the same real-world events. A separate skill worth naming: repost-chain pivoting, recognizing "this is a repost, not original" and following its outbound link to the source.
Beyond a single account, look at the network surrounding a subject: who they interact with, follow, get tagged by. The most practically useful idea here is second-order (third-party) leakage: a friend tags them at a venue, a colleague mentions their employer by name, an event page bills them as a speaker. A privacy-conscious user can carefully clean their own timeline and still be exposed by an organization's public page or a friend's tag. You don't control what your own network posts about you.
For larger networks (bot networks, coordinated campaigns), manual analysis doesn't scale. Social Network Analysis (SNA) uses graph theory: nodes are accounts, edges are connections (mutual follows, say), weight is the strength of interaction (how many retweets/replies), degree is how many connections a node has, and centrality is how important a node is relative to the whole network. Free tools like Gephi or Maltego Community Edition make these calculations visually accessible.
In 2019, a Bellingcat investigation analyzed several days of the hashtags #FreeWestPapua and #WestPapua, building a dataset of usernames, retweets, likes, and timestamps. It then used Gephi to isolate a cluster of accounts showing an obviously non-organic, coordinated pattern pushing a pro-Indonesian government narrative. This was a good demonstration of how visualization surfaces patterns manual reading would miss entirely.
Content verification increasingly overlaps with OSINT, because the spread of false information is now part of the investigation itself. Three distinct categories:
| Type | Definition | Example |
|---|---|---|
| Misinformation | Misleading/wrong, not knowingly deceptive | Someone shares a false health article in good faith |
| Disinformation | Misleading/wrong, deliberately | A deliberately false narrative crafted to stir conflict |
| Malinformation | Based in reality but deployed to cause harm | A true fact presented out of context to incite hatred |
Starting questions when you're looking at suspect content: Who posted it? What's the motive? Who benefits? Can you trace the origin? What behavior is the account showing? The process looks a lot like reverse engineering: find the earliest instance of the claim (screenshot every step, since the original can be deleted), use archiving tools, and keep a spreadsheet of URLs to spot patterns.
We cover AI-generated content and deepfake detection specifically (a fast-moving area) in Chapter 13, alongside the older manipulation-detection techniques from Chapter 6.
Bots mimic human behavior but leave traces: generic or cartoon profile photos (avoiding the risk of reverse-image exposure that comes with using a real photo), usernames with an obviously automated pattern and near-identical creation dates, an unusually high posting rate paired with very few followers, and mutual networking/amplification within the group. One well-documented historical example of scale: Russia's Internet Research Agency ran thousands of fake Twitter accounts around the 2016 US election period, collectively generating millions of tweets and a large aggregate follower count. It's a reminder of how fast a modern disinformation campaign can scale.
Comment threads and buried forum replies are often richer than headline posts, since people are less guarded in conversation, assuming no one scrolls that far. Also: what the browser renders is only part of the page. View-source / DevTools reveals alt text, aria-labels, image title attributes, and HTML comments that never appear visually: a photo that looks uninformative can carry alt-text the uploader wrote describing exactly what it shows.
Old material matters precisely because it's less curated and often predates a subject's later privacy-awareness. A core SOCMINT judgment call: distinguishing where someone lives from where they merely visited. Travel posts mark where someone was at that moment, not where they live: wording ("getting out of the city," "back home") usually separates a trip from a home base, and repetition across time settles the question (home shows up again and again, a trip shows up once). A discipline worth adopting: for every candidate location finding, write down the anti-finding too — what would have to be true if this were a trip rather than a residence. Then check which version the body of posts actually supports better. Watch too for correlation vs. coincidence: handle reuse and superficial matches breed false positives, since someone else entirely could be using the same username, or a namesake. Establish identity through several independent signals before attributing anything.
Keep observation ("posts at 8:05am"), inference ("commutes ~8:00-8:30am"), and conclusion ("works near the city center") clearly separated. Any finding is only as solid as its weakest verified link. In practice, keep a separate record for each element you're extracting: residential area, transit line, work area, weekly free day, with fields for the finding, its evidence (post, platform, text), the verification, and a confidence level. One record per element stops findings from blurring into each other and effectively pre-writes the report.
Every clue you uncover is an OPSEC mistake on the subject's part. Geotags left on expose a workplace. Consistent posting times expose a routine. Handle reuse across platforms enables cross-platform correlation. Third-party exposure lets an event page reveal what the subject withheld. Unscrubbed metadata leaks alt-text or affiliation. Partial OPSEC isn't OPSEC: removing GPS data while leaving alt-text, geotags, and regular timing in place still yields a full profile. Full treatment in Chapter 16.
Everything so far in this chapter assumes a platform run by one company with one central database: the model that Facebook, X, and Instagram all share. A growing share of activity now happens on federated networks, where that assumption breaks down. Mastodon (part of the broader "fediverse," built on a protocol called ActivityPub) and Bluesky (built on a different protocol, the AT Protocol) both spread accounts and content across many independently-run servers instead of one company's servers.
What this actually changes for an investigator: a Mastodon account doesn't live on "Mastodon" the way a Twitter account lives on Twitter. It lives on one specific, independently-operated instance (a server run by some individual, community, or organization, each with its own domain, its own moderation policy, and its own rules about what's publicly visible to non-members). Two instances can behave completely differently. One might index publicly to search engines, another might not. One might federate openly with most other instances, another might block large parts of the network entirely. Before you can even start investigating a Mastodon account, you need to know which instance it's on and what that instance's specific visibility settings actually allow you to see. Bluesky is more centralized in practice today than the fediverse model suggests, but is built to allow the same kind of portability and independent hosting over time.
The good news: the core techniques from earlier in this chapter still apply directly: handle reuse, profile-photo correlation, and posting-pattern analysis (§5.2, §5.4) work the same way on a federated account as on a centralized one. What's new is that cross-platform correlation now has to happen at both the account level and the instance level. The same person might maintain a consistent handle across a Mastodon instance and Bluesky, and confirming it's genuinely one person still relies on the same corroborating-signal discipline from §5.4 — just applied across a wider, less centralized set of platforms.
Discord servers, in-game economies (Roblox's in-game trading and marketplace features are a well-documented example), and community features built into platforms like Steam were never designed as public broadcast platforms the way Twitter or Instagram were. That design difference is exactly what makes them a genuinely harder OSINT domain, not just another platform to add to a checklist.
The core difficulty is discoverability. A public tweet is indexed, searchable, and visible to anyone by default. A Discord server's content typically isn't: it's often invite-only, and joining the server is frequently the only way to observe anything happening inside it at all. That single fact changes the investigative posture significantly. The passive-vs-active distinction from §5.1 becomes sharper here, because simply gaining visibility into the space at all may already cross from passive observation into something closer to active engagement. The same OPSEC exposure risks apply as in Chapter 16 — a joining account is an account that can be noticed, tracked, and potentially unmasked by the community it joined.
This matters because these semi-private spaces have real investigative relevance on both ends of a spectrum. They've been used for organizing communities that deliberately seek spaces with weaker outside visibility, including extremist community-building. Separately, game-adjacent chat has raised well-documented child-safety concerns around grooming, precisely because a game environment gives an adult a plausible, low-suspicion reason to be interacting with much younger users in real time. Neither of these is a reason to treat gaming platforms as inherently suspicious. The overwhelming majority of activity on them is exactly what it looks like, ordinary social and recreational use, but it's why they show up in investigative contexts more than their "just a game" reputation suggests.
Practically: server and channel structure varies enormously. A server might have dozens of channels with different membership and visibility rules, and member lists may or may not be visible depending on server settings. Message history retention is controlled entirely by the server's own configuration (some retain everything, some auto-delete after a set period) rather than by any platform-wide policy you can rely on. Bot logs (automated moderation or activity-tracking bots that many servers run) can be a useful secondary source when they exist, but their presence, configuration, and what they retain is entirely server-specific.
The bot-spotting signals in §5.6, namely generic photos, automated-looking usernames, and a high posting rate with few followers, catch individual fake accounts. They miss the more consequential pattern: real influence and disinformation operations increasingly work at the network level, where a mix of automated and genuinely human-operated accounts coordinate in ways that no single account's profile would ever reveal on its own. Platform trust & safety teams and researchers use the term coordinated inauthentic behavior (CIB) for exactly this: behavior that looks organic account-by-account but reveals itself as coordinated once you look at the accounts together.
Four indicators are worth checking, though none is conclusive alone. Timing coordination: supposedly-unrelated accounts posting near-identical content within minutes of each other, repeatedly, across many separate incidents — occasional coincidence happens, but a repeated pattern doesn't. Account-creation clustering: a batch of accounts all created within a tight window, later activated together for a single campaign. Template-like phrasing: near-identical wording reused across accounts that otherwise present as unrelated individuals. Network structure: an amplification pattern (who retweets, shares, or reposts whom) that looks artificially symmetric or centrally directed compared to how organic sharing normally spreads through a real social network.
This is precisely where the Social Network Analysis tooling from §5.5 earns its place. Gephi (or Maltego Community Edition) is how you actually see the network-level pattern that individual-account review can't surface — the same way the Bellingcat #FreeWestPapua case in §5.5 visualized a coordinated cluster that wouldn't have been obvious from reading posts one at a time. The discipline to hold onto: a single account showing bot-like signals is a lead, not proof of coordination. CIB is inherently a network-level judgment, and the finding only becomes solid once you can show the pattern across multiple accounts, not describe one suspicious-looking profile.
Everything else in this chapter is written around investigating a specific, already-identified target. A different, complementary skill is standing monitoring: tracking an evolving topic, entity, or threat continuously over time, rather than researching one subject once and moving on. A simple five-step shape covers most monitoring work. Define what you're actually monitoring — specific entities, keywords, or narratives, narrow enough to stay useful, but not so narrow that you miss obvious variants and misspellings. Prefer structured collection over manually re-running the same search: an API or RSS feed you can poll on a schedule beats remembering to re-search by hand, and scales far better. Deduplicate and filter before you ever sit down to review results, since the same claim reposted two hundred times is one data point, not two hundred. Set explicit thresholds for what actually warrants a closer look, so routine background noise doesn't consume the review time you need for genuine signal. And preserve as you go: take a timestamped snapshot the moment you see something, not a mental note for later, because monitored content is disproportionately likely to be exactly the kind of content that gets deleted once its author notices it drew attention.
Chapter 8 §8.7 already works through one fully concrete version of this exact five-step shape, applied specifically to dark-web monitoring (SOCKS5h connection handling, periodic circuit rotation, tiered keyword severity, context-snippet capture). It's worth reading as a detailed worked example of how this general pattern looks once you commit it to an actual implementation in one specific domain, rather than repeating that detail here.
For exercise 1, don't over-claim. A rough active-hours window from a handful of data points is a reasonable finding, but a precise timezone from three posts is over-reaching. For exercise 3, resist the urge to call an account a "bot" outright: write the finding as "shows N indicators consistent with automation" and name them, which is exactly the confidence-calibrated language a real report needs (see Chapter 17). For exercise 5, the most common mistake is defining the monitoring scope too broadly ("anything about topic X"). A workable monitoring target is specific enough that your threshold step in §5.14 actually has something concrete to filter against.