
Weaponizing Exposed Data — data indexing, secondary abuse and the emerging risk of soft data pivots. SOURCE · LAB-1
Executive Summary

The shift from bulk-dump extortion to structured, indexed and priced data. SOURCE · LAB-1
Threat actors are changing what they do with stolen data. Whereas the old incumbent model of extortion relied on bulk publication of data and the victim’s fear of what an archive might contain, a growing number of groups now analyze, index and structure exfiltrated material before it is published or sold. Removing the cost of weaponization converts stolen data from an undifferentiated liability into a searchable, priced and targetable asset with significant impact on victims, third-parties and individuals.
The effect of this is multifold including elevated impact during negotiation, increased likelihood of sale through tranching and pricing, elevated personal exposure for named individuals, and extended impact to third parties via chained attacks. As of July 2026, at least 10 groups are performing data analytics on stolen data relying on commodity models and freely available tooling as well as bespoke infrastructure.
Some groups have significantly increased their capability and usage of data analytics creating bespoke breach incident reports against victims using human analysis, Artificial Intelligence (AI), Machine Learning (ML) and file path analysis on stolen data to create detailed reporting, structured findings and a highlight of the most sensitive data taken. The extreme details and understanding of stolen data enhance the threat posed by such groups via downstream targeting and higher reputational, regulatory and legal exposure.
This report examines the effect of this trend and is relevant to a technical and executive audience.
Key Highlights
- Throughout Q1–Q2 2026, a rapidly growing number of Ransomware-as-a-Service (RaaS) and pure data-extortion groups have moved beyond unstructured bulk dumps toward deliberately analyzing and indexing the data they steal. This post-exfiltration analysis has become a resourced capability, engineered to amplify extortion leverage, raise secondary-sale value and deepen downstream harm to victims and their stakeholders. Threat actors including FulcrumSec, Settra, ShinyHunters, Sentap, The Gentlemen, Everest, Anubis, Titan and Leak Bazaar are all engaging with some form of data analytics at a rising pace showcasing the increasing threat posed by this trend.
- Lab-1 has observed three principal methodologies employed by groups conducting post-exfiltration data analysis throughout 2026; human analysis, AI, ML & LLM driven analysis and basic scripted extraction. While these are separate, the boundaries between them are increasingly porous, and the more capable groups combine several methodologies enabling high-fidelity extortion attacks.
- Reorganizing stolen data around criminal utility and secondary abuse removes the ‘weaponization tax’ - the level of time, effort and skill required to utilize stolen data - which increases the threat posed by data extortion by driving a range of adverse effects including regulatory punishments, secondary weaponization and elevated reputational and personal impact. The burden of locating damaging material within an undifferentiated archive that once shielded victims is increasingly performed by the adversary and, in many cases, published openly dictating that a breach an organization might have contained as an IT incident is increasingly converted into regulatory, criminal, reputational, civil and competitive risk.
- Lab-1 assesses with high confidence that data analytics compounds the impact of a breach across three dimensions. During negotiation it replaces the abstract threat of a bulk dump with a structured inventory of exactly what was taken, categorized by regulatory sensitivity and financial exposure, which materially increases the pressure to pay. It raises the probability of selling the data by dividing a dataset into separately priced tranches that lower the entry price and widen the buyer pool. It also elevates the risk of secondary abuse as the analytical work needed to surface credentials, client relationships and pressure points has been completed on the buyer’s behalf.
- Hard data pivots like credentials, API keys, tokens, certificates and other technical artifacts are identified and exploited to grant direct, authenticated access. Soft data pivots exploit non-technical material: invoices, organizational charts, client lists, project documentation and internal correspondence that is being weaponized to enable context driven social-engineering attacks and downstream targeting. Lab-1 assesses with high confidence that while hard pivots remain the fastest and most reliable route to chained attacks, they are also anticipated by defenders, who revoke them post-breach. The soft data pivot exploits the unaddressed assumption that non-technical data is harmless, and represents one of the most underestimated risks associated with data disclosure events.
TABLE 01 — Dump Disclosure Model vs Data Analytics Dumps
| Dimension | Incumbent Bulk-Dump Model | Emerging Data Analysis Model | Risk |
|---|---|---|---|
| Buyer pool | Small; requires effort or database tooling to extract value from raw dumps. | Expanded; ML-cleaned, category-segmented products are consumable by non-technical buyers. | ELEVATED |
| Criminal Utility | Capped at the willingness-to-pay of a single bulk buyer. | Lower purchase price via tranched data increases sale chance. Price discrimination across buyer classes raises total monetizable value. | SIGNIFICANTLY ELEVATED |
| Downstream attack risks | Low; most data is never used because it is unknown and may not link to any buyer objective. | High; curated tranches reach the specific buyer class optimized to weaponize them. | SIGNIFICANTLY ELEVATED |
| Exposure window | Single public-disclosure event; often stabilizes within weeks. | Continuous resale tail over months or years; chronic re-surfacing through new buyers. | ELEVATED |
| Coercive leverage | Primarily confidentiality and reputational. | Curated regulatory evidence (OFAC/SDN, shadow accounting, unissued financials). | SIGNIFICANTLY ELEVATED |
| Victim observability | Dump is typically public; scope can be assessed directly. | Tranches sold exclusively may never surface; harm appears downstream without a forensic trail. Or data may be made much more publicly available due to easier download and curated insights. | ELEVATED |
| Defender cost | Focused on a single disclosure event and its immediate fallout. | Scales with tranche count, buyer diversity, and ongoing monitoring of brokerage platforms. | ELEVATED |
Section 01 — Background & Trends
Data exfiltration has become a near-universal component of ransomware operations with an estimated 100% of the currently most active (top 10) ransomware groups performing data exfiltration as part of ransomware operations, or focusing entirely on pure data extortion without ransomware. Coveware’s Q4 2025 incident data recorded exfiltration in 94% of ransomware cases [Coveware], while BlackFog’s Q3 2025 analysis of 449 dark web victim listings placed the average exfiltration volume at 527.65 GB per victim [BlackFog]. The Sophos State of Ransomware 2025 report places the figure at 75% of attacks involving exfiltration prior to encryption, and records encryption itself falling to 49% of attacks, down from 66% in 2024 and the lowest rate in the survey’s five-year history [Sophos].
Arctic Wolf’s 2026 threat report records data-extortion-only attacks, where no encryption is deployed at all, rising elevenfold between November 2024 and November 2025, from 2% to 22% of its incident-response caseload [Arctic Wolf, via Cybersecurity Dive], with Sophos recording the same shift at smaller scale as extortion-only attacks rose to 10% from 3% in 2024 [Sophos]. Together these figures reflect a broader operational shift toward data-only extortion models.
In extension of this, Lab-1 assesses with high confidence that a significant shift in the ransomware and data extortion group threatscape is ongoing with a rising volume of groups focusing on data analytics and indexation of stolen data after exfiltration has occurred. Throughout 2025 and Q1 and Q2 of 2026, the post-exfiltration phase of the attack lifecycle has evolved from being merely an unstructured bulk dumping to deliberate, resourced, data analytics designed to amplify extortion leverage, raise secondary sale value and elevate downstream harm and malicious utility. This is a noteworthy shift as groups increasingly focus on data extortion over ransomware, making it paramount that victims now understand their own exposure or third-party exposure, more in depth and quicker, than threat-actors do.
Figure 01 — The evolution of extortion: from undifferentiated bulk dumps toward structured, indexed and analytically weaponized data. SOURCE · LAB-1
Section 02 — When Threat-Actors Analyze Your Data
Lab-1 and RS-ISAC first observed data indexing in the wake of Ashley Madison when threat-actors built an indexing website so that spouses or colleagues could see if anyone they knew was on there. In a ransomware context, Lab-1’s Dark-web Research Team first observed data indexing in 2022 when ALPHV-ng and Industrial Spy began deploying searchable leak-site interfaces that allowed visitors to query stolen datasets by keyword or file type. These early efforts were limited in scale and sophistication, restricted to basic search functions across narrow data categories, and were assumed to be prohibitively expensive to maintain.
Figure 02 — The evolution of threat-actor data analytics over time. SOURCE · LAB-1
By Q1 2026, Lab-1 tracks multiple groups performing significant data indexing and analytics at scale. Analysis of multiple case studies of victims affected by ransomware and data extortion groups indicates a near-tenfold increase in groups doing data analytics instead of the older pure dump model. At least 10 currently conduct some form of structured data analysis (up from 1 in 2024), as evidenced by the groups’ own statements and the way exfiltrated material is presented, weaponized, indexed and sold. Below is a table highlighting some current and past groups engaging with stolen data analytics.
TABLE 02 — Groups Engaging in Stolen-Data Analytics
| Group | Type | Indexing Method | Evidence & Observable Activity |
|---|---|---|---|
| FulcrumSec | Pure Data Extortion | Unknown AI models, AI Agents + Human Analysis | Deployed a team of AI agents against 1.3 TB of Novo Nordisk data (700,717 files) to identify and highlight high-impact material. Human-driven tradecraft observed at Arup (7-month credential spidering, structured DLS presentation) and MCO (isolated 61 OFAC-sensitive emails from 395,211 email archives). Mixture of automated AI analysis and human curation. |
| The Gentlemen | RaaS | AI, LLMs + Human Analysis | Analyzed exfiltrated Adaptivist data to map downstream client relationships, identify Turkish targets and extract contextual intelligence for follow-on attacks. Created backdoor Okta accounts within hours of pivoting. Operational pattern consistent with dedicated human analysis rather than automated tooling. |
| ShinyHunters | Pure Data Extortion | Unknown AI + Human Analysis | ShinyHunters acknowledged the use of unknown AI tools to analyze the data stolen from NAIC. The AI analysis was used to drive severity ratings by ShinyHunters and the assessed impact. |
| Titan | RaaS | AI + NLP Models | Structures stolen datasets around high-impact categories designed to maximize victim pressure and buyer value. Lab-1 Dark-web Research Team assesses dedicated team involvement based on the quality and categorization of indexed outputs. |
| Everest | Pure Data Extortion | Unknown AI | Showed likely signs of basic AI driven data analysis in their “PT Dharma archive” which they claimed contains over 157,000 detailed file records with 133 GB of data. |
| Sentap | IAB & Pure Data Extortion | AI / LLMs | Lab-1 assesses Sentap actively use LLMs and other AI-driven data analytic tools to examine stolen data on dark-web forums. Demonstrated indexed outputs derived from stolen datasets. |
| Leak Bazaar (SnowTeam) | Indexing-as-a-Service | ML + Human Analysis | Advertised pipeline combines ML-assisted text analysis, automated debris removal, database reverse engineering (SQL, SAP, Oracle), ERP parsing and human analyst validation. 16-category taxonomy across 6 languages. At least 8 clients including The Gentlemen RaaS. |
| Settra | Pure Data Extortion | LLM + unknown AI | Lab-1 assesses with moderate confidence that the newly emerged data extortion group Settra utilize LLMs and AI models to understand stolen data. Evidenced against the group’s breach of Conduril. |
| Xpl0itrs | IAB & Pure Data Extortion | Unknown AI | Xpl0itrs was observed performing data analytics on data stolen from Rapidfort, allegedly in collaboration with TeamPCP on TierOne forum. Veracity unconfirmed. |
| Anubis | RaaS | Dedicated Team | Lab-1 assesses dedicated team involvement based on the structured presentation of indexed victim data on the group’s DLS. |
| World Leaks | RaaS (rebranding of Hunters International) | Basic Scripting | Early indexing attempts observed in 2022 with searchable DLS interfaces. Limited data categories and basic search functionality. |
| ALPHV (BlackCat) | RaaS · Defunct | Basic Scripting | First observed deploying searchable leak-site interfaces in 2022. Basic keyword and file-type search across stolen datasets. Low sophistication relative to current groups. Early mover in the indexing space. |
| Industrial Spy | Marketplace · Defunct | Basic Scripting | Operated a marketplace model with basic search and category functions from 2022. Limited scale. Platform defunct. |
Examples of Observed Threat-Actor Activity
Lab-1 has mapped and summarized 8 groups or threat actors in Figure 03 below assessed to be engaged with data analytics. A number of these groups and threat actors are examined in more detail in the following case studies.
Figure 03 — Groups and threat actors assessed by Lab-1 to be performing data analytics on stolen data. SOURCE · LAB-1
FulcrumSec
Lab-1’s Dark-web Research Team assesses with moderate confidence that FulcrumSec represents one of the most operationally mature examples of threat-actor data analytics observed by Lab-1 to date. In the group’s breach of Arup, disclosed on 13 May 2026, FulcrumSec gained initial access through a hardcoded GitHub personal access token discovered on a forgotten sandbox subdomain granting read access to approximately 10,000 private repositories across the Arup GitHub organization. FulcrumSec spidered credentials from repository to repository, pivoting across 29 Azure Blob Storage accounts, 44 AWS S3 buckets containing 4.4 million files, AWS Redshift and RDS databases, a Neo4j graph database, SharePoint sites and a GCP project with production payment gateway credentials.

EXHIBIT 1 — FulcrumSec data-leak-site post structuring the Arup breach around regulatory and reputational pressure categories. SOURCE · FULCRUMSEC DLS
In the aftermath of the breach, Lab-1 finds that FulcrumSec did not merely publish the stolen data but rather spent 6 months analyzing their stolen loot, with the group structuring and categorizing the material on its data leak site, highlighting not only their attack path but also how the stolen data poses direct risk to Arup clients and to Arup itself, surfacing direct data, links, relationships and information affecting those clients (see Exhibit 1).
FulcrumSec applied the same methodology to its breach of Novo Nordisk, allegedly deploying a team of AI agents against 1.3 terabytes of exfiltrated data across 700,717 files to identify and highlight high-impact material including 11,500 patient records across named clinical trials, 41,144 proprietary compound structures, 32 private AI models and five undisclosed drug programs. FulcrumSec targeted specific informational pressure points including alleged drug trials, intellectual property, and sensitive healthcare data. In the group’s breach of MyComplianceOffice (MCO), FulcrumSec’s analytics uncovered 61 specific emails, out of a data-set of nearly 400,000 email archive files and 200 GB of data overall, which allegedly showcase MCO helping clients bypass international sanctions. The group then structured the exfiltrated data to spotlight the 61 emails touching OFAC-sanctioned entities, thereby having an elevated impact compared to a mass data dump.
FulcrumSec’s consistent pattern across Arup, Novo Nordisk and MCO likely represents a deliberate operational model with the group investing significant resources post-exfiltration in understanding, structuring and weaponizing stolen data before engaging victims or publishing.
‘Titan’ Ransomware-as-a-Service
The newly emerged RaaS group ‘Titan’ has demonstrated comparable investment in post-exfiltration analysis, structuring stolen datasets around high-impact categories designed to maximize victim pressure and buyer value.
The group brands its RaaS operation as built around a proprietary AI engine running on dedicated infrastructure operating as a “zero-touch” AI capability which, once connected to a victim environment, autonomously classifies documents by sensitivity, maps relationships between counterparties and offshore entities, identifies the specific files whose exposure would cause maximum damage, calculates regulatory penalties and prepares publication queues without manual intervention and calculates an “optimal” ransom demand calibrated to the victim’s financial position (see Exhibit 2). The advertised engine capabilities include:

EXHIBIT 2 — TITAN’s data-leak site advertising its on-premises AI analysis engine. SOURCE · TITAN DLS
- Automated document classification: categorization by sensitivity across financial records, legal documents, PII, trade secrets, intellectual property and internal correspondence.
- Tax evasion detection: dataset-wide scanning reaching back to 1998 for hidden transactions, undeclared revenue, falsified invoices and systematic avoidance patterns.
- Entity relationship mapping: surfacing connections between counterparties, offshore entities, shell companies and affiliated persons.
- Critical asset identification: pinpointing the specific files whose exposure would inflict maximum reputational, legal and financial damage.
- Automated legal impact assessment: jurisdiction-specific regulatory analysis tied to the GDPR, Singapore’s PDPA, Italy’s D.Lgs. 231/2001, California’s CCPA/CPRA and a claimed 50-plus further frameworks.
- Regulatory Notification Packages: auto-generation of ready-to-send notification letters addressed to tax authorities, data protection agencies, financial intelligence units, law enforcement and national media outlets.
- Ransom Calculation Engine: a demand calibrated to the victim’s financial position, revenue, cash flow and total regulatory exposure, set high enough to be meaningful but low enough to remain payable.
Lab-1 assesses with high confidence that this shift focuses output on pre-packaged automated distribution to the regulators and journalists best placed to inflict secondary harm on the victim.
ShinyHunters
On June 11 ShinyHunters - a notorious and prolific data extortionist group - confirmed the use of AI to triage stolen data. The group listed the National Association of Insurance Commissioners (NAIC) claiming exfiltration of roughly 3.1TB (see Exhibit 3).

EXHIBIT 3 — ShinyHunters’ data-leak listing for the National Association of Insurance Commissioners (NAIC) breach. SOURCE · SHINYHUNTERS DLS
ShinyHunters subsequently revised its description of the dataset, attributing its initial overstatement of the contents to an “AI-generated misinterpretation of the underlying data” that was corrected after human review. This is the first confirmed case of ShinyHunters using AI to drive data-analytics observed in the wild. The episode is especially relevant in light of NAIC’s rebuttal, which disputed the scope of the claim and stated that core regulatory systems and personally identifiable information were not accessed, with only statutory financial reports, investment credit-rating data, and outdated logs and configuration files taken.
The notion that ShinyHunters struggled to accurately characterize the data and that the resulting claim became a contested narrative between adversary and victim reinforces a central finding in this report wherein exposure is increasingly defined by whoever indexes the data first; the capacity to rapidly and accurately establish what was actually taken has become decisive for the perpetrator and victim alike.
The Gentlemen
Lab-1 has observed The Gentlemen RaaS group focusing on data-indexing capabilities that converts bulk exfiltrated files into a structured, searchable intelligence product. Against at least two victims the group has published bespoke “Breach Incident Reports”; in one observed case, a 54-page report on the Colombian energy major Ecopetrol, hosted as a dedicated page on the group’s own leak infrastructure. The report classifies the entire stolen file tree into eleven business categories and assigns each a severity rating, a narrative analysis, extracted content signals, and descriptive tags identifying operational hubs, cross-folder data duplication, and the time span of the material.
The data is a curation rather than a summary and quantifies the haul in a complete file-type inventory, more than 300,000 files and in excess of a terabyte of data, including on the order of 109,000 PDFs, 85,000 spreadsheets and 19,000 presentations and a “Critical Findings” section that hand-selects individual high-value files by path and size, each with a one-line note explaining its significance. The effort further pre-computes the victim’s regulatory exposure, citing specific breach-notification regimes and deadlines including specific GDPR Articles, Colombian data-protection law and upstream-sector obligations as explicit leverage. A summary of the exposure can be seen in Exhibits 4 and 5 below taken from The Gentlemen DLS and compiled in the graphics by Lab-1.

EXHIBIT 4 — The Gentlemen showcasing heavy data analytics on stolen data. SOURCE · THE GENTLEMEN DLS

EXHIBIT 5 — The Gentlemen showcasing heavy data analytics on stolen data. SOURCE · THE GENTLEMEN DLS
The Gentlemen is especially noteworthy as Lab-1 has observed members of The Gentlemen RaaS operation actively discussing the use of AI tools to facilitate post-exfiltration data analysis. In communications leaked from the group, the actor “Protagor”, assessed by Lab-1 to be affiliated with The Gentlemen, described a methodology in which exfiltrated data of high importance is provided to an AI system for review and categorization.

EXHIBIT 6 — Leaked messages from ‘Protagor’ proposing rented GPU on vast.ai to run “uncensored AI” against stolen data. SOURCE · The Gentlemen Leaked Data via Checkpoint[1]
Settra
Lab-1 assesses with moderate confidence that the newly emerged data-extortion group Settra are also engaging in structured post-exfiltration data analysis, an assessment based on the structure of its public leak-site postings. Notably, Settra appear to be one of the groups performing the most granular analysis on their stolen loot.
In Settra’s listing for the Portuguese construction firm Conduril, advertised at approximately 196.6 million records and around 163 GB, Settra organizes the archive around impact and audience. The post is structured with a “Who this matters to” section addressed separately to regulators (Portugal’s CNPD, the securities regulator CMVM, and Angola’s AGT), to employees, to international sponsors and clients (the World Bank, AFD, EIB, US EXIM, and MCC), and to competitors and consortium partners.

EXHIBIT 7 — Settra’s data-leak-site publication on Conduril, structured by stakeholder audience and surfacing special-category personal, medical and financial data. SOURCE · SETTRA DLS

EXHIBIT 8 — Settra’s data-leak-site publication on Conduril, structured by stakeholder audience and surfacing special-category personal, medical and financial data. SOURCE · SETTRA DLS
Lab-1 finds that this stakeholder-based layout replaces the victim’s original file structure with a map of who is exposed and how. The group surfaces the highest-leverage material it holds, including special-category medical data of extreme sensitivity, financial data and a corporate internet-banking login and password held in plain text. Settra further quantifies the exposure in a highly specific document-category table listing as seen in Exhibits 7 and 8.
Xpl0itrs
On 22 July 2026, xpl0itrs, an Initial Access Broker (IAB) and pure data extortionist operating on TierOne, advertised 569GB of data attributed to RapidFort, allegedly exfiltrated during the CanisterWorm campaign conducted alongside TeamPCP, priced at USD 40,000 (see Exhibit 9). Lab-1 assesses with high confidence that the listing evidences deliberate data indexation as the seller enumerates 140,061 files across 48 S3 buckets and publishes a bucket-level manifest giving size, file count and a functional description of each store, distinguishing the development hardening pipeline from the production hardening bucket and the Azure nightly debug environment.

EXHIBIT 9 — Xpl0itrs Showcasing Data Analytics on Stolen RapidFort Data. SOURCE · TIERONE FORUM
Of particular note is the seller’s consolidated extraction and listing of high-potency material like credentials discovered across the full corpus into a single deduplicated inventory, itemizing six AWS credential pairs, two kubeconfig files carrying service principal authentication to AKS clusters, GitLab PostgreSQL and replication credentials, a full Azure storage account key, RSA private keys and EC2 instance credentials. Lab-1 notes that each entry is typed by the access it confers, which showcase a value assessment on top of the extraction itself.
The credential inventory is consistent with the output of commodity secret-scanning utilities run across the corpus, bucket sizes and file counts are recoverable directly from S3 enumeration, and the descriptive text is terse and informal, lacking the narrative construction, regulatory mapping and financial modeling that characterize the AI-assisted operators examined earlier in this report. Lab-1 finds that the case demonstrates that indexation delivers material advantage, including identifying RapidFort customers and drawing attention to the presence of Department of Defense data.
Data Analysis Methods Observed Across Groups
Assessing Capability from Disclosure and Behavior
Establishing what analytical capability a group deploys is one of the harder judgments in this research as disclosure is inconsistent. Some groups describe their tooling in detail while others disclose no insights, referring only to “advanced tooling” or “new capabilities” without naming models, infrastructure or method. Where a group does not declare its capability, Lab-1 assesses it from the observable characteristics of the output and from the group’s behavior across forums. Five indicators carry most of the weight:
- Specificity of the claim. Named models, infrastructure, throughput figures and cost detail raise confidence. Generic capability language carries little evidential weight on its own, and claims drawn from a group’s own marketing are treated as unverified.
- Structure of the published output. Examination of whether material is categorized, indexed, priced by tranche, cross-referenced between documents and mapped to specific regulatory regimes. Output organized around buyer demand and victim pressure, in place of the victim’s original file structure, indicates a deliberate processing step between exfiltration and publication.
- The volume-to-precision ratio. The relationship between dataset size and the precision of what is surfaced is the single strongest indicator. Isolating key data from large files, quickly, is difficult to reconcile with manual review alone.
- Comparison against the pre-AI baseline. Where possible, benchmark current output against the same group’s earlier activity and against the bulk-dump norm that preceded this shift. A step change in the structure, speed or precision of publication points to new tooling.
- Language and error signatures. Narrativized write-ups, annotation conventions typical of large language models, and triage of material in languages the operator does not appear to read all suggest model assistance. Self-correction is especially instructive as seen when ShinyHunters attributed its own overstatement of the NAIC dataset to an “AI-generated misinterpretation” corrected only after human review, which evidenced the capability more firmly than any marketing claim.
Figure 04: Confidence Assessment of Analytics Capabilities
Observed Methodologies
Three principal methodologies have been observed employed by groups conducting post-exfiltration data analysis: human analysis, AI & LLM driven analysis and basic, scripted extraction. While these are described separately, the boundaries between them are increasingly porous, and the more capable groups combine several methodologies with each converging on the same objective; converting raw stolen data into structured, actionable information to pressure victims, elevate impact or increase chance of secondary sale. These are summarised in Figure 05 below.
Figure 05 — The data-analytics capability spectrum, from basic scripted extraction through AI and ML pipelines to dedicated human analysis. SOURCE · LAB-1
Human Analysis
Dedicated human analyst teams represent the first category. The approach was pioneered by Anubis with Lab-1 assessing that Anubis maintains dedicated analysts within its organization who craft exposure pieces resembling the work of investigative journalists rather than that of ransomware operatives, narrativizing stolen data into structured, public-facing accounts designed to maximize reputational pressure on the victim. FulcrumSec has demonstrated a comparable model, evidenced by the group’s operation against Arup and possibly also Novo Nordisk and MCO (see Exhibits 10 and 11).

EXHIBIT 10 — FulcrumSec’s data-leak-site write-up of the Arup breach & Novo Nordisk AI models — narrated like investigative journalism. SOURCE · FULCRUMSEC DLS

EXHIBIT 11 — FulcrumSec’s data-leak-site write-up of the Arup breach & Novo Nordisk AI models — narrated like investigative journalism. SOURCE · FULCRUMSEC DLS
In a notable case observed by Lab-1, an actor operating under the moniker “AUDITOR” advertised data analysis services to ransomware operators on a prominent underground forum, claiming extensive experience auditing for the Big Four (KPMG, PwC, Deloitte and EY). The actor offers auditor capabilities to ransomware groups enabling them to identify specific, valuable data before an attack is launched, analyzing the target and determining which files will specifically guarantee payment, including accounting records, internal correspondence, tender documents and incriminating information on key individuals (see Exhibits 12 to 14).

EXHIBIT 12 — ‘AUDITOR’ advertises Big-Four-grade audit services to ransomware operators on an underground forum. SOURCE · TIERONE FORUM

EXHIBIT 13 — ‘AUDITOR’ advertises Big-Four-grade audit services to ransomware operators on an underground forum. SOURCE · TIERONE FORUM

EXHIBIT 14 — ‘AUDITOR’ advertises Big-Four-grade audit services to ransomware operators on an underground forum. SOURCE · TIERONE FORUM
The actor distinguished this approach from standard ransomware operations, noting that:
“standard ransomware has a 50% success rate if backups are available. My approach guarantees leverage, even if the system is rolled back. The victim pays not for decryption, but for silence.”
After exfiltration, the actor conducts an audit of the stolen material and prepares an initial company message outlining the consequences of refusing to cooperate. By 27 April 2026, the actor posted a follow-up claiming to “successfully increase payout conversion rates.” AUDITOR further highlights that partnering with him is beneficial before an attack, reducing the amount of data needing to be taken due to better insights, or after to better find pressure-points based on already exfiltrated information.
Lab-1 assesses with moderate confidence that this represents a professionalization of the post-exfiltration analysis function, where individuals with legitimate financial audit training are offering their skills as a service to threat actors who lack the domain expertise to identify high-leverage material within stolen datasets.
Artificial Intelligence & Machine Learning
Artificial intelligence (AI) and machine learning (ML) constitute the second category and represent the current operational frontier, in which large language models (LLMs) or other custom AI tooling are used to scan, categorize and extract value from data volumes. Sentap exemplifies the accessible end of this spectrum, relying on commodity models and commercially available tooling, what they call “advanced tooling”, rather than bespoke infrastructure, to index bulk hauls and surface high-value material, thereby converting a raw dataset into structured leverage against the victim and, in turn, against its clients and suppliers. Lab-1 assesses with moderate confidence that multiple groups are currently attempting to use AI for data analytics.
The Gentlemen illustrate how readily this capability can be stood up outside the commercial AI ecosystem altogether. Notably, Lab-1 has observed group-affiliated discussion centered on provisioning “uncensored” open-weight LLMs on the GPU-rental marketplace vast.ai in order to process sensitive stolen data without the refusal behaviors built into services such as ChatGPT or Claude.
At the most capable end of AI usage sits the Titan ransomware group, which claims GPU-accelerated inference on dedicated AMD EPYC infrastructure, a claimed throughput of up to 700GB of mixed corporate documents in under an hour, and custom Natural Language Processing (NLP) models trained on a corpus exceeding ten million corporate documents, financial statements, tax filings, employment contracts and legal texts across multiple jurisdictions and languages. While these figures are drawn entirely from Titan’s own claims and remain independently unverified, Lab-1 nonetheless finds significance in the operating model focused on a fully autonomous pipeline. These pre-packaged analytical outputs with automated distribution appear designed to offer regulators and journalists the ability to conduct investigations and to inflict additional harm (see Exhibit 15).

EXHIBIT 15 — TITAN describing the technical architecture of its AI analysis engine. SOURCE · TITAN DLS
This typology is also observed by FulcrumSec who in addition to human analysis also claimed to have deployed a pipeline of AI agents against the 1.3 terabytes and 700,717 files exfiltrated, paired with an AI-driven quality-assurance layer that reviewed and validated the agents’ output before it was structured for publication[2] in a recent attack. Lab-1 assesses this dual use of autonomous AI, for both the analysis itself and the quality control of that analysis, as a significant development that demonstrates how even operations built on strong human tradecraft are now embedding agentic AI into their workflows in order to multiply the output of their human analysts.
Script Extractions
The third category relies on traditional approaches, including GREP-based searches, scripted extraction routines and database query tools. While the least sophisticated, it remains effective and widely used. This methodology predates AI and was primarily adopted by ALPHV-ng, Industrial Spy and Hunters International enabling quick identification of invoices, organization charts or payment documents. Lab-1 assesses that AI has subsumed this approach, whereby keyword and regular-expression filtering may serve as the first-pass triage before LLM indexation (see Exhibit 16).

EXHIBIT 16 — ALPHV’s searchable leak-site interface returning keyword results for ‘invoice’ across indexed stolen data. SOURCE · ALPHV DLS
Lab-1 assesses that the convergence of these methods, combined with the increasing accessibility of AI tools capable of processing unstructured enterprise data, will continue to lower the barrier to entry for data indexing. Groups that could not deploy dedicated analyst teams in 2024 can now deploy commodity LLMs against stolen datasets at negligible cost while capabilities such as TITAN’s AI models promise to place an entire analytical pipeline, from classification through legal impact assessment, ransom calculation and pre-generated regulatory notifications, into the hands of affiliates with no analytical capability of their own.
Figure 06 — The three principal post-exfiltration data-analysis methodologies: human analysis, AI and ML, and basic scripted extraction. SOURCE · LAB-1
Capabilities are Built to Target Enterprise Data Repositories
Some threat actors build their capabilities specifically to target the systems in which high-value corporate data is concentrated, rather than collecting indiscriminately across a compromised network. Leak Bazaar is the clearest example with an advertised pipeline engineered around the canonical enterprise data stores, performing automated database reverse engineering against SQL, SAP and Oracle systems and enterprise resource planning (ERP) parsing to analyze table structures and extract the most monetizable records, financial transactions, payroll and contractor data, before rendering them into clean Excel or CSV output ready for immediate use (see Exhibit 17).
This showcases tooling purpose-built for the formats and platforms in which structured enterprise data resides enabling actors to align their tooling to the specific platforms where the most valuable corporate data is held. The effect is to make data indexing faster, more reliable and more repeatable, lowering the weaponization tax still further and signaling an increasingly sophisticated, platform-aware understanding of enterprise data environments among dark-web operators.

EXHIBIT 17 — SnowTeam’s full ‘Leak Bazaar’ advertisement on TierOne, 25 March 2026, describing its ML pipeline and DBMS/ERP reverse-engineering capability. SOURCE · TIERONE FORUM
Section 03 — Spotlight — Leak Bazaar: a Data-Analysis-as-a-Service Platform
While the above examples showcase in-house capabilities for threat-actor usage of AI and human analyst to develop data analytical capabilities, Lab-1 has increasingly also seen the capability offered as a service on dark-web forums whereby threat-actors offer RaaS-agnostic indexing as a processing layer, for a fee thereby extending the availability of data indexing and elevating the threat posed by the rapidly expanding trend.
For example, on 25 March 2026, the threat actor “Snow” / “SnowTeam” advertised on the TierOne (T1) forum, a high-end Russian-language cybercrime forum specializing in ransomware, announcing a new criminal service “Leak Bazaar.” The offering is a focused attempt to professionalize the post-exfiltration market, repositioning raw stolen corporate data as a refined, category-segmented intelligence product sold to distinct buyer classes rather than dumped wholesale for extortion leverage alone. Leak Bazaar offers a corporate data exchange built on ML-powered dump analysis, database management system (DBMS) reverse engineering, and ransomware negotiation support. The listing positions the platform as a destination for operators who have already exfiltrated large volumes of corporate material, but who lack the tooling, personnel, or commercial infrastructure to convert those dumps into monetizable intelligence products with in-depth data segmentation and analysis (see Exhibit 18).

EXHIBIT 18 — Pinned advertisement posted by Snow on TierOne, 25 March 2026, introducing Leak Bazaar as a closed corporate data exchange. SOURCE · TIERONE FORUM
Snow explicitly positions the service as infrastructure for converting “dead weight into real profit” with the threat actor pitching to sellers that “No one needs to buy 5 TB of raw data. If you, as a technical specialist, are only interested in the company’s R&D, you purchase only this segment.”
Leak Bazaar accepts corporate data originating from organizations with annual revenue above USD ten million and imposes a minimum submission threshold of 100 GB, with a stated preference for one terabyte or more. The service only handles unpublished, non-Russian, non-CIS material, commercially valuable or actionable Western corporate intelligence rather than recycled, previously leaked datasets.
According to the advertisement the platform’s server cluster performs NLP analysis of text arrays using filtering algorithms designed by a professional mathematician, automated DBMS and ERP reverse engineering against SQL, SAP and Oracle exports, and debris removal that filters out OS backups, system files and non-valuable content followed by human analysis conducting final validation before being segmented into buyer-oriented product categories.
Data Taxonomy and Multilingual Targeting Expand the Exposure Surface
Lab-1 notes that the advertised pipeline explicitly reorganizes information into market-driven categories rather than preserving the victim’s original file structure. The operator published a formal sixteen-category search taxonomy together with a six-language targeting list. Lab-1 assesses with moderate confidence that the taxonomy is engineered as a buyer-navigation as well as a victim-organization surface, meaning the categories correspond to the downstream use cases a buyer is shopping for while also teasing the exposed victim data in detail from the original dump (Exhibit 19).

EXHIBIT 19 — Operator publication of the sixteen-category search taxonomy and six-language targeting list on TierOne, 3 April 2026. SOURCE · TIERONE FORUM
Lab-1 has mapped each category to its most probable buyer type and associated risk mechanism in the table below. The mapping is analytical and not restrictive: in practice, individual tranches are likely to be consumed by more than one buyer class, and the taxonomy itself enables the cross-correlation attacks described in the Analysis section.
TABLE 03 — Category, Buyer & Risk Mechanism
| Advertised Category | Probable Buyer Archetype | Risks |
|---|---|---|
| Quarterly Reports | Insider-trading operators, short-selling networks | Pre-announcement market positioning; securities exposure |
| Financial Records | Fraud rings, forensic accounting leverage | Payment redirection, AML and tax exposure |
| Mergers & Acquisitions | Traders, competing bidders, nation-state adjacent | Deal leakage, bidder frontrunning, regulatory disclosure risk |
| Share Buybacks | Market-manipulation actors | Pre-announcement trading; Reg FD exposure |
| Dividend Records | Insider-trading operators | Pre-announcement positioning |
| Divestment Plans | Competitors, activist shorts | Strategic disclosure risk |
| Cost of Production | Competitors, commodity traders | Margin disclosure, bidding intelligence |
| Suppliers & Buyers | Business email compromise (BEC) operators, competitors | Supply-chain pivot, fraudulent invoicing |
| Guidance Documents | Competitors, regulators-adjacent actors | Strategic intelligence, governance risk |
| Research Reports | Competitors, industrial-espionage operators | R&D theft, patent-race acceleration |
| Sanctions Lists | Ransomware negotiators (coercion), state-adjacent buyers | OFAC/SDN exposure as extortion leverage |
| Regulatory Violations | Ransomware negotiators, plaintiff firms, regulators-adjacent | Enforcement leverage, civil litigation exposure |
| Security Reports | Follow-on intrusion operators, competing attackers | Vulnerability roadmap for re-compromise |
| Insurance Documents | Fraud rings, extortion operators | Coverage-gap exploitation, policy fraud |
| Confidential Data | Broad-buyer class | Default coercive leverage category |
| Organization Structure | BEC operators, social engineers, insider-recruitment actors | Pretexting, lateral social engineering |
The six-language targeting list (English, German, Spanish, French, Turkish, and Italian) expands the victim aperture beyond the English-speaking frameworks and victims.
Actively Operating
On 19 April 2026, Snow published a sample archive listing on the TierOne thread which showcased downloadable product lots. The listing references six archives totaling approximately 15.75 GB with filenames whose suffixes align directly to the 16-category taxonomy published on 3 April 2026 (Exhibit 20).

EXHIBIT 20 — Archive listing published by the operator on the TierOne Leak Bazaar thread, 19 April 2026. SOURCE · TIERONE FORUM
Lab-1 assesses with moderate confidence that the archive listing represents the earliest observable operational delivery of the Leak Bazaar pipeline and showcase regulatory exposure while the bulk financial record remains available as a separately-priced tranche.
TABLE 04 — Sample Archive Lots · 19 Apr 2026
| Lot (Category Suffix) | Size | Mapped Category |
|---|---|---|
_com_sanctions.zip | 92.26 KB | Sanctions Lists |
_organization.zip | 3.09 GB | Organization Structure |
_violations.zip | 1.41 GB | Regulatory Violations |
_finance.zip | 10.71 GB | Financial Records |
_buyback.zip | 251.55 MB | Share Buybacks |
_research_reports.zip | 190.64 MB | Research Reports |
Lab-1 observes a sequence of endorsements and feedback from vetted members of TierOne on the Leak Bazaar thread between 25 March and 21 July 2026 that collectively raise actor-reliability while also confirming the service as being operational, therefore raising the threat posed by Snow’s offering (see Exhibits 21 and 22).

EXHIBIT 21 — TierOne members endorse the Leak Bazaar listing. SOURCE · TIERONE FORUM

EXHIBIT 22 — TierOne members endorse the Leak Bazaar listing; member “beebeek3” endorses on 1 April 2026. SOURCE · TIERONE FORUM
Section 04 — How Analyzed Data Elevates Enterprise Risks
Lab-1 assesses that the emergence of structured data indexing, whether conducted in-house or outsourced, changes the risk profile of every data breach in which it is applied. When threat actors reorganize stolen data around demand, criminal value and victim pressure, they remove the weaponization tax. Before the emergence of stolen data analytics, this tax functioned as a barrier to the weaponization of stolen data and protected direct and third-party victims by the technical difficulty of downloading the data sets and the prohibitively hard burden of deriving insights from stolen data sets. Even when data was stolen and published, the cost of parsing terabytes of unstructured corporate files was prohibitive for most buyers, journalists and regulators. Indexed data eliminates that buffer entirely. Threat actors such as Sentap, FulcrumSec and Leak Bazaar absorb the weaponization tax on behalf of their buyers and reorganize stolen data around downstream value which Lab-1 assesses compounds impact across four primary dimensions: elevated extortion impact during negotiations, higher chance of a sale of the data through lower prices and higher criminal utility, increased personal exposure and much higher chance of third-party impact.
Figure 07 — The four dimensions across which indexed data compounds breach impact: extortion leverage, sale probability, personal exposure and third-party risk. SOURCE · LAB-1

EXHIBIT 23 — Archive listing published by the operator of Leak Bazaar on the TierOne Leak Bazaar thread, 27 April 2026. SOURCE · TIERONE FORUM
Elevated Extortion Impact
Lab-1 assesses that data indexing materially increases the leverage available to threat actors during extortion negotiations. When a group can present a victim with a structured inventory of exactly what has been stolen, categorized by regulatory sensitivity, financial exposure and reputational impact, the extortion dynamic shifts from abstract threat to concrete evidence whereby, instead of the probability that something damaging might surface from a bulk dump, the victim is confronted with a curated catalog of sensitive material, presented with the specificity of an internal audit. This was observed by The Gentlemen in the group’s two latest “Breach Incident Reports” whereby the group found and presented critical findings mapped to regulatory frameworks while also presenting data relevant for class action lawsuits or other high impact findings (see Exhibits 4 and 5).

EXHIBIT 24 — FulcrumSec’s DLS post detailing the MCO breach and its 61 OFAC-sensitive emails. SOURCE · FULCRUMSEC DLS
FulcrumSec’s operations demonstrate this pattern across Arup, MyComplianceOffice (MCO) (see Exhibit 24) and Novo Nordisk. In each case the structured presentation was itself a tool, demonstrating to the victim that the threat actor understood the data as well as its own internal teams (see Exhibit 25).

EXHIBIT 25 — FulcrumSec data-leak-site post detailing the exfiltration of Novo Nordisk research data. SOURCE · FULCRUMSEC DLS
In a listing observed by Lab-1 for data stolen from GreenHills Ventures, a venture capital firm, Sentap categorized 34 GB of exfiltrated material into Legal Documents (NDAs, Series A and B term sheets, shareholder agreements and accredited investor forms), Financial Data (capitalization tables, balance sheets and three-year financial forecasts) and Organizational Data (organizational charts, management team profiles and compensation reports). This was aimed at highlighting the data most likely to have a significant impact on GreenHills Ventures based on their industry and what data they consider valuable and sensitive.
The Gentlemen’s analytics efforts also raise the threat to victims and to third parties. It sharply increases extortion leverage, by presenting the most damaging material — crown-jewel intellectual property, commercial terms, and executive and workforce data — at the front of a professional-looking report making the cost of non-payment immediate and legible to a victim’s leadership and readily citable by journalists, competitors and regulators. Because the data is heavily curated, the harm extends well beyond the primary victim to contractor pricing, third-party contracts, partner records and individuals’ personal and medical data are all surfaced, and the exposure of physical-security plans and operational engineering data for critical energy infrastructure introduces potential safety and physical-security consequences that extend past data privacy.
Lab-1 assesses that this level of post-exfiltration processing is designed to signal to both victims and buyers that the data has been verified, structured and is ready for use. By indexing stolen data around regulatory sensitivity, financial exposure and individual culpability, they pre-assemble the precise evidentiary material that regulators, prosecutors, litigators and competitors require to initiate class-action lawsuits or other fines.
Case Study: Titan AI Engine
Lab-1 has obtained the analytical output associated with multiple TITAN RaaS listings (see Exhibits 26–29). The output takes the form of a financial exposure assessment addressed directly to the victim, structured across regulatory fines, direct fraud liability, breach notification and remediation, class action settlement exposure, competitive damage and reputational cost, closing with a summary table and a single headline estimate of USD 1.5 million to 2.5 million.
EXHIBIT 26 — TITAN’s AI Generated and Analysed Loss Assessments. SOURCE · TITAN DLS
Across the incidents the typology the effort produces is granular and moves well beyond generic data categories. From the leak it separates personally identifiable information, comprising names and addresses, from financial account data, internal audit communications and proprietary plan documents, and then resolves that material down to named document classes including trustee certification statements, trial balances, summaries of net trust assets and participant loan default reports. It surfaces classification markings applied by the victim itself, isolating material labeled “Trade Secret” and “Internal Only”, and identifies commercially sensitive subject matter through hits on gross margin, research and development, acquisition, investor and shareholder content. The engine additionally reconstructs context that is not stated in any single file, identifying the plan sponsor as a corporate entity distinct from the plan, estimating a participant population of 50 to 150 from the file listing alone, and inferring the presence of social security and employer identification numbers where these are not explicitly labeled.
Lab-1 finds that the more significant capability is the engine’s ability to match identified files to specific enterprise risk. The output goes beyond a file inventory and attaches each category of document to the liability channel it activates, associating participant data and plan financial statements with fiduciary exposure under ERISA, driver’s license, passport and financial account combinations with the statutory damages provisions of California privacy law, bank and account numbers with direct fraud and fraudulent tax filing, and trade secret and margin material with competitive loss. It further establishes jurisdiction from corporate location, identifying the victim’s California presence as the trigger for state privacy exposure and flagging GDPR applicability as conditional on the presence of EU participants. Lab-1 assesses with high confidence that this represents a meaningful advance on bulk publication, since the actor presents the victim with a pre-built account of which specific documents create which specific legal consequence, in a form requiring no interpretation by the reader.
For example, the identification of an internal cybersecurity policy document is one of the most consequential findings within the output as the policy document sets out the controls, standards and incident response procedures an organization states it operates, and its presence in an exfiltrated corpus allows the actor to measure the victim against its own documented commitments thereby calibrating ransom demands against the victim’s own documented policies. Where a breach has occurred despite those commitments, the document becomes direct evidence of a gap between stated and actual control, which is material to regulatory examination and to any subsequent negligence or fiduciary claim. Lab-1 further assesses that the document carries direct negotiating value beyond its evidentiary weight, since it is likely to disclose the victim’s incident response playbook, escalation thresholds, notification timelines, insurance requirements and any standing position on ransom payment. An actor holding that material can time pressure against the victim’s regulatory notification deadlines and can anticipate the victim’s response before it is made.
Lab-1 assesses with high confidence that the monetary figures within the output are calibrated for negotiating effect and represent a “worst case scenario” rather than an accurate assessment of loss. The regulatory attributions are noteworthy but do not appear entirely accurate, and several of the totals rest on inferred rather than confirmed record counts. This accuracy shortfall does little to reduce operational risk, since the effect is to remove the victim’s ability to define the loss narrative during negotiation. A secondary risk follows, in that the actor’s own estimate may be repeated in public reporting and become the reference figure for the incident.

EXHIBIT 27 — TITAN AI-engine financial-exposure output report. SOURCE · TITAN DLS

EXHIBIT 28 — TITAN AI-engine financial-exposure output report. SOURCE · TITAN DLS

EXHIBIT 29 — TITAN AI-engine financial-exposure output report. SOURCE · TITAN DLS
Higher Chance of Sale
Lab-1 assesses that data indexing expands the buyer pool for stolen data by lowering the unit price and structuring material around specific buyer profiles and criminal value. The traditional bulk-dump model constrained the data resale market to technically capable actors willing to invest in parsing raw data and willing to pay high sums for an archive whose contents were uncertain. Indexing separates a stolen dataset into multiple saleable tranches, each priced and packaged for a defined buyer category. This drives up the overall chance of a sale and thereby extends the risk associated with leaked data.
“I found on the Leak Bazaar all categories I needed. I need it for my project. Everything I need was included. I’m satisfied.” — B17, a satisfied Leak Bazaar buyer, TierOne

EXHIBIT 30 — Buyer praising the Leak Bazaar sale model. SOURCE · TIERONE FORUM
This is evidenced by Leak Bazaar which breaks with the traditional model operated by RaaS groups, historically optimized for coercion against “the next victim” rather than an actual sale of stolen data. Leak Bazaar’s advertised pipeline instead focuses on usability of stolen data by transforming the dump into a set of segmented, buyer-specific products aligned to distinct criminal use cases including ransomware-negotiation leverage, social-engineering pretexting, and follow-on intrusion operations including ransomware operators, BEC operators, fraud rings, insider-trading networks, industrial-espionage buyers, and regulatory-leverage actors (see Exhibit 31).

EXHIBIT 31 — SnowTeam’s pitch: segmented, per-category sale replaces the bulk dump. SOURCE · TIERONE FORUM
Sentap’s listing for data stolen from Intuitive Machines, a US space technology company further illustrates this model. Sentap divided 70 GB of exfiltrated material into nine separately priced categories. The total asking price for the complete dataset was $100,000 exclusive or $50,000 non-exclusive whereas the tranche model means a buyer interested only in cost-of-production data can acquire it for $7,500. A buyer interested only in potential regulatory violations can acquire that tranche for $500. This dictates that material that would never have found a buyer within a bulk archive now has a market at an accessible price point.
EXHIBIT 32 — Sentap’s data-leak listing dividing 70 GB of Intuitive Machines data into nine separately priced tranches. SOURCE · SENTAP DLS
As evidenced by FulcrumSec’s attack against Novo Nordisk, the indexation of stolen data enables FulcrumSec to sell various tranches of data to various buyers. The group claims to be seeking data purchasers for their stolen data specifically highlighting value for competitors like Eli Lilly or “a well-funded biotech in a jurisdiction with flexible attitudes toward IP provenance” (see Exhibit 33).

EXHIBIT 33 — FulcrumSec data-leak-site post offering the stolen Novo Nordisk data for sale to competitors. SOURCE · FULCRUMSEC DLS
Personal Exposure & Data Aggregation
Lab-1 has observed groups using indexed data to build targeted narrative packages designed to create the impression that the C-suite is complicit in the breach, either through failure to secure the data or unwillingness to contain the breach after it occurred. These packages target individuals by name, surface compensation data, personal communications and decision-making records, and are designed to fracture internal cohesion between upper management, employees, customers and third parties. In extreme cases observed by Lab-1, this has led to executive doxxing, career-targeted attacks and physical harm risks including swatting. The objective is to create pressure vectors that extend beyond the organization’s ability to contain through conventional incident response, forcing payment by making the personal cost of non-payment unbearable for the individuals named.
Lab-1 assesses that data indexing amplifies this risk category specifically because it enables threat actors to surface the exact material, executive compensation, security audit failures, compliance shortfalls, undisclosed financial exposures, that creates the most damaging narrative when presented publicly. In a bulk dump, this material exists but is functionally hidden. In an indexed dataset, it is the first thing a buyer or journalist sees.
In addition, when breaches affect customer data and private individuals, Lab-1 assesses with high confidence that data aggregation and identity threats pose an elevated risk. The personal exposure created by an indexed dataset is not bounded by the originating breach, but compounds when the material is cross-referenced against records already circulating from prior incidents. Indexing accelerates this aggregation by enabling the abuse of individuals’ PII across sources until the composite profile substantially exceeds what any single source contained. Lab-1 notes that large-scale breaches corroborate this pattern, whereby freshly stolen records are routinely merged with data from earlier incidents to build progressively richer victim profiles, and that the underlying data retains its utility indefinitely once it reaches criminal markets.
Lab-1 assesses with high confidence that aggregation of this kind increases impersonation risk, and that partial or seemingly low-sensitivity fields are sufficient to enable it. Increasingly identifiers such as the final digits of a card or account number, a date of birth, or a home address are enough to satisfy the knowledge-based checks that banks, telecommunications providers, online retailers, and internal help desks use to authenticate a customer. This can assist account takeover attacks, unauthorized transactions against payment methods already held on file by a merchant, and social engineering of support staff. This is amplified as breach material can be used to manufacture fraudulent identity documents enabling identity-proofing schemes and lower the threshold for establishing credit or accounts in a victim’s name thereby harming customers and affected entities. Lab-1 assesses with high confidence that the data can be paired with dark-web AI services and tools including vocal and video deepfake capabilities, synthetic identity generation and elevated social-engineering via AI call centers. Voice cloning and synthetic media will likely feature with increasing frequency in these campaigns with the aggregated breach data supplying the contextual detail that renders the synthetic artifact credible.
Third-party Targeting and Impact
Lab-1 assesses with high confidence that data indexing increasingly facilitates chained breaches, in which data stolen from one organization is used to enrich, drive or directly enable the compromise of a second organization within the supply-chain. Lab-1 notes that in this scenario, third-party risk is not merely data exposed from a third-party, but a potential direct compromise via the third-party insights.
Figure 08 — Chained attacks: how soft-data and hard-data pivots convert a single breach into a multi-victim campaign. SOURCE · LAB-1
The Gentlemen’s compromise of Adaptivist, a UK-based Atlassian consultancy, represents a case study of this pattern. In April 2026, The Gentlemen breached Adaptivist’s infrastructure and exfiltrated material allegedly including 20,000+ JIRA Legal Services tickets, 33,000 legal documents, 24,547 Confluence pages comprising over 100 GB of internal documentation, 484,220 HubSpot CRM customer contacts, Nexus repositories containing Helm charts with production secrets and over 3 TB of Docker images, and ScriptRunner’s complete proprietary codebase and licensing system. Critically, The Gentlemen did not treat this as a standalone extortion event but instead weaponized data from Adaptivist to target their clients. According to leaked internal communications seen by Lab-1, data exfiltrated from Adaptivist was reused both before and during a subsequent attack against Arcelik, a Turkish consumer electronics manufacturer with approximately $11.8 billion in annual revenue for whom Adaptivist had provided consultancy services.
Initial access to Arcelik was obtained independently via a vulnerable FortiGate VPN appliance. However, the Adaptivist data provided operational intelligence that materially enhanced the Turkish campaign. zeta88, The Gentlemen’s administrator, explicitly referenced data from the UK consultancy breach to cross-reference and enrich information about the Turkish target, including an internal Transfer/Migration Document that described project work the consultancy had performed for Arcelik. Internal communications observed by Lab-1 show Protagor, one of nine named core operators within The Gentlemen’s affiliate structure, analyzing the credential landscape of Adaptivist’s downstream clients, identifying three primary entry points including Okta credentials, VPN login and password combinations, and FortiGate credentials. zeta88 subsequently created a backdoor Okta service account within the Arcelik environment during the operation.
The Gentlemen then published both Adaptivist and Arcelik on their data leak site, explicitly labeling Adaptivist as the “access broker” for the Turkish attack. This served a dual purpose that of punishing the UK consultancy, which the operators described as “a very bad company,” and increasing pressure on the Turkish company by promising to demonstrate exactly how the breach originated, encouraging Arcelik to pursue legal action against Adaptivist. Lab-1 assesses that this case demonstrates how indexed and analyzed data converts a single breach into a multi-victim campaign.
The TeamPCP supply chain campaign illustrates the same principle through a different mechanism. In early 2026, TeamPCP exploited credentials that were not fully revoked following a prior, distinct security incident at Aqua Security. This residual access enabled TeamPCP, on 19 March 2026, to force-push malicious commits to 76 of 77 version tags in Aqua Security’s Trivy GitHub Action repositories, as well as all 7 tags in the setup-trivy repository. Credentials harvested from Trivy were subsequently used to access Cisco’s build pipelines, clone over 300 internal GitHub repositories including source code for unreleased AI products, and move laterally into a small number of Cisco’s AWS accounts. ShinyHunters then claimed the Cisco data, publishing an extortion post on 31 March 2026 with a 3 April deadline, alleging theft of over 3 million Salesforce records alongside the cloned repositories and AWS assets. In a parallel chain, the same Trivy compromise enabled TeamPCP to breach the European Commission’s Europa web hosting platform, with CERT-EU confirming on 2 April 2026 that 340 GB (91.7 GB compressed) had been exfiltrated from the compromised AWS account, including data pertaining to up to 71 clients of the Europa web hosting service: 42 internal clients of the European Commission and at least 29 other Union entities. Lab-1 notes that in both cases, the originating compromise was residual access from an incompletely remediated security incident at a developer security tool vendor. The blast radius extended to a Fortune 500 technology company, a supranational governing body and 71 affiliated organizations, none of which had a direct security relationship with the original vulnerability.
Section 05 — Outlook
The shift toward structured data analysis by threat groups elevates enterprise risk as exposed data now poses a materially higher risk to victim organizations than at any prior point. Previously, the difficulty of making sense of stolen data provided a de facto buffer, even when credentials, tokens or sensitive documents sat within a leaked archive, the likelihood of a secondary actor finding, extracting and operationalizing that specific material was low. While the data existed, it was functionally inaccessible within the noise. That buffer has lowered with groups now actively mining exfiltrated datasets for material that enables follow-on attacks, and they are doing so across two distinct categories that Lab-1 classifies as hard data pivots and soft data pivots.
Looking ahead, Lab-1 assesses with high confidence that this threat will intensify as analytical capability matures and as service providers such as Leak Bazaar industrialize the indexation, structuring and resale of stolen data. Moreover, as efficacy of bulk data dumping decreases, more groups may pivot to analytical disclosure attacks. The most significant consequence is the expanding exposure of a data analytics attack acting as a continuous malicious event, in which the disclosure no longer marks the end of risk.
Hard Data Pivots: Credentials, Tokens and Technical Artifacts
The traditional secondary risk from a breach has always centered on what Lab-1 terms hard data: credentials, API keys, access tokens, cloud secrets, SSL certificates, SSH private keys and other technical artifacts that provide direct, authenticated access to systems. This risk is well understood and has driven post-breach response protocols for years: rotate credentials, revoke tokens, invalidate keys. But the scale and speed at which hard data is now being extracted and operationalized is elevated as evidenced by TeamPCP against Aqua Security and their downstream victims.
Lab-1 assesses with high confidence that while hard data pivots will remain the fastest and most reliable mechanism for chained attacks it is also the strongest focus for defenders, reducing the ability for threat-actors to weaponize the pivot.
Soft Data Pivots: An Overlooked Risk
Increasingly, however, the consequential secondary attacks are originating from what Lab-1 terms soft data: non-technical information such as invoices, organizational charts, client lists, project documentation, internal communications, supplier contracts and commercial correspondence. This material has traditionally been classified as low-risk in breach impact assessments because it does not provide direct system access. Lab-1 assesses with high confidence that this classification is no longer adequate.
The Gentlemen’s breach of Adaptivist is the clearest observed case study of a soft data pivot in operation as seen with The Gentlemen weaponizing an internal Transfer/Migration Document that described project work Adaptivist had performed for Arcelik. This is the soft data pivot which weaponizes information that the majority of organizations would not classify as sensitive into attack-enabling intelligence, for example, an invoice stolen from a breach can be weaponized to enable Vendor Email Compromise or Business Email Compromise. Other soft data can enrich context driven social-engineering campaigns or an email chain can enable an attacker to craft an email referencing the exact purchase order number, addressed to the correct billing contact, requesting a payment redirect that matches the expected amount. The same risk applies to organizational charts that reveal reporting lines and delegation of authority, to project documentation that exposes vendor relationships and contract values, to quarterly reports that disclose financial positions and commercial sensitivities, and to internal communications that reveal the language patterns and decision-making processes of key individuals. While there is no technical indicator, the soft-data pivot has enabled high fidelity attacks via contextual details that are accurate. Lab-1 notes that each of these data categories is routinely exfiltrated in current breaches and each is increasingly being indexed, structured and sold to buyers who understand how to operationalize it.
Section 06 — Mitigations: Rapid Understanding of Exposed Data Becomes Increasingly Important
The rapid evolution in threat-actor data analysis, and the emergence of service providers such as Leak Bazaar, makes it imperative that organizations understand the scope of their exposure and the risk it carries for stakeholders faster than an adversary can index it.
Know and Reduce What Can Be Indexed
- Maintain a current data map covering what the organization holds, where it resides, which system of record owns it, and which records are regulated (special-category personal data, PHI, payment data, trade secrets, source code). The map is what allows exposure to be scoped within hours.
- Enforce retention and deletion schedules against the repositories that actually accumulate data: email archives, shared drives, ticketing systems, CRM exports and backup snapshots. Every archive kept past its retention period is indexable inventory.
- Minimize and tokenize special-category data at rest, pseudonymizing medical, criminal-record and biometric fields wherever the business process does not require the raw value.
- Remove secrets from data stores. Scan repositories, Confluence and SharePoint pages, configuration files and ticket attachments for hard-coded credentials, tokens and keys, and move them into a managed vault.
- Seed high-value repositories with canary documents and honeytokens. A file or credential that alerts when opened, queried or resold gives notice that a dataset is being indexed and can identify the buyer.
Harden the Repositories Actors Build Tooling Against
- Apply least privilege and phishing-resistant MFA to the systems this report identifies as targets: SAP and Oracle ERP, Salesforce and other CRM, SQL databases, Snowflake and BigQuery warehouses, and cloud object storage including AWS S3 and Azure Blob.
- Restrict and alert on bulk export. Cap or require approval for mass record exports, full-table dumps and large report generation, and log the identity behind every bulk query.
- Inventory and restrict OAuth and connected applications across SaaS platforms, revoking unused third-party grants on a fixed cycle.
- Rotate service-account and integration credentials on schedule, and prefer short-lived tokens over long-lived personal access tokens.
Detect Mass Collection Before Publication
- Alert on collection behavior as well as on exfiltration: unusual volumes of read or query activity, enumeration of storage buckets, recursive directory crawling and repository cloning at scale.
- Monitor for staging activity, including the creation of large archives and the transfer tooling commonly used for bulk movement.
- Baseline and alert on egress volume by user, service account and destination, with particular attention to cloud-to-cloud transfers that bypass the corporate perimeter.
Contain Hard-Data Pivots Completely
- Treat every credential, API key, access token, SSH key, certificate and cloud secret present in an exfiltrated dataset as compromised, and rotate on that assumption without waiting for evidence of use.
- Revoke active sessions and refresh tokens alongside password resets, because credential rotation on its own leaves live sessions intact.
- Re-scan for residual access after remediation. The TeamPCP campaign in this report originated in credentials that were never fully revoked following an earlier, separate incident.
Close the Soft-Data Pivot Gap
- Require out-of-band verification, using previously known contact details, for any change to bank details, payment instructions or payee records, however convincing the referenced invoice or purchase-order number appears.
- Brief finance, accounts-payable and procurement teams that adversaries now hold genuine invoices, purchase orders, contract values and billing contacts, so contextual accuracy no longer serves as evidence of legitimacy.
- Alert on inbox-rule creation and mail-forwarding changes, and enforce DMARC, DKIM and SPF with external-sender tagging.
- Notify named suppliers, clients and partners whose commercial data appears in the exfiltrated set, since they are the population the soft-data pivot targets next.
Extend Exposure Management Beyond the Perimeter
- Monitor for the organization’s own data inside third-party breaches, covering consultancies, managed service providers and consortium partners. The Adaptivist case shows the downstream victim frequently has no visibility of the originating breach.
- Require by contract that suppliers notify the organization of a breach within a defined period, with sufficient scope detail to determine which of its records were involved.
- Maintain a register of which partners hold which categories of the organization’s data, so a partner breach can be scoped immediately.
Sustain Exposure Assessment Across the Resale Tail
- Treat exposure assessment as a continuing process. A tranche-based sales model converts a single breach into a resale tail that persists for months or years, and refined derivatives may surface long after the original dump.
- Monitor marketplaces and leak sites for derivative listings and category-aligned lot naming. Where a lot suffix matching a known taxonomy appears beside a partially redacted victim prefix, treat it as an early indicator of a prior, potentially undisclosed breach entering the brokerage market.
- Re-run exposure analysis whenever new tranches, samples or derivative reports appear, and update notifications accordingly.
Protect Named Individuals
- Identify executives and staff named in exfiltrated material and brief them directly on targeted narrative packages, doxxing and physical-harm risks including swatting.
- Reduce the personal data available for enrichment by removing executive and family information from data-broker and people-search services, and review home-address exposure.
- Extend monitoring to the personal email and social accounts of named individuals for the duration of the extortion window.
Prepare to Characterize and Communicate Exposure
- Assume the adversary has access to analysis-as-a-service. AUDITOR and Leak Bazaar make triage available to operators with no in-house capability, so scope planning should assume rapid, accurate identification of regulated and high-leverage records regardless of the actor’s own sophistication.
- Build the capability to answer “exactly what was taken” as a standing function, with tooling and access rehearsed in advance and ready before an incident begins.
- Engage legal and regulatory counsel at the point of discovery, given the criminal, civil and regulatory dimensions this report documents.
- Prepare disclosure to regulators, customers and partners on the assumption that the adversary will publish a structured, categorized account first. Communicating exposure accurately and proactively, before the adversary defines the narrative, is now a core part of incident response.