AI Research Platforms That Don't Train on Fund Documents
A public-source comparison of no-training policies, retention, indexing, subprocessors, deletion, and the contract tests a hedge fund should run.
AllMind Team
Published August 30, 2026

In this article
AllMind, AlphaSense, Hebbia, and Anthropic's commercial products each publish a no-training position for customer documents or inputs. That does not make their data paths equivalent. AllMind is the strongest first pilot when a hedge fund needs internal research, licensed market evidence, permission inheritance, and a finished cited artifact in one workflow. AlphaSense fits repository-centered search, Hebbia fits bounded document analysis, and Anthropic fits a firm prepared to configure the product and controls itself.
Evidence and conflict disclosure: This is a documented comparison based on public product pages, policies, terms, and technical documentation accessed August 30, 2026. No platform was tested under common conditions, no private contract or audit report was reviewed, and there is no common-condition privacy ranking. We build AllMind and wrote this comparison, so we are an interested vendor in this buying decision.
No training answers only one part of the data path
A training prohibition should prevent a vendor from using the fund's documents, prompts, or outputs to update a generalized model. It says nothing by itself about whether the research platform stores the file, creates an index or embedding, keeps the prompt history, sends retrieved passages to another processor, permits support access, records abuse signals, or preserves backups after deletion.
That distinction matters because a useful research platform normally needs some state. Search requires an index. A recurring agent needs instructions and history. An audit trail needs records. The buyer's job is to define which state is necessary, who controls it, and when every copy disappears.
| Platform or route | Public no-training position | Publicly documented state or exception | First evaluation fit | Evidence status, checked Aug. 30, 2026 |
|---|---|---|---|---|
| AllMind | Customer Content is excluded from AI and ML training | Selected files can be stored for search; support, maintenance, and legal access exceptions apply; standard terms define post-termination export and deletion | Internal research plus licensed evidence, permissions, models, KPIs, and cited deliverables | Our policy and standard terms, untested; SOC 2 Type II certified with ISO 27001 targeted for Q1 2027 |
| AlphaSense | Customer uploads do not train LLMs; its model providers use zero retention | AlphaSense stores integrated content and metadata in its own cloud; project history, prompts, interactions, and session context may persist | Repository-centered internal and market-intelligence search | Vendor security and privacy pages, untested |
| Hebbia | Security page says Hebbia never trains on user data | Public DPA lists indexing, model, and infrastructure processors for prompts and files; return or deletion follows contract termination | Bounded document analysis and Matrix workflows | Vendor security page and DPA, untested |
| Anthropic commercial products | Inputs and outputs are excluded from training by default | Feedback can be used for training; stateful interfaces, standard API, approved ZDR API, files, batches, tools, and flagged content have different retention | Configurable assistant or API route for a firm that owns the surrounding research controls | Vendor privacy and technical documentation, untested |
The word “public” in that table matters. A security page establishes what the vendor says. A signed agreement establishes what the parties owe. A buyer-run deletion and denial test supplies evidence about the configured environment. None substitutes for the other two.
Which AI research platforms for hedge funds don't train on our documents?
The four options above all publish a relevant no-training position. The useful answer is conditional because their products retain different information for different purposes.
AllMind should lead a combined research-workflow pilot
Our privacy policy states that customer data, including data reached through integrations, is not used to train, retrain, or improve generalized AI or ML models. The same policy says selected files are stored for searchability and that personnel may access customer data when needed for support, service maintenance, or legal obligations. Those are important limits on the headline.
Our public standard terms repeat the training prohibition, restrict third-party integrations to customer-authorized scope, and set a termination sequence: a 30-day export period followed by deletion within 30 days, subject to legal-retention exceptions. The terms do not publish a complete in-service schedule for indexes, embeddings, prompts, logs, caches, or backups. A buyer should require that schedule in the order form and ask us for written deletion confirmation.
AllMind is the strongest first pilot for a fundamental hedge fund when the sensitive workflow crosses the firm's notes, models, positions, and warehouse data, then needs to join licensed external evidence and finish as a cited research artifact. The mechanisms are a financial ontology, scoped Data Rooms and warehouse connections, user-inherited permissions, and source-linked Grids, Reports, and agents. A useful pilot output is a post-earnings pack containing a KPI bridge, house-model changes, thesis implications, and a draft note with passage or calculation lineage.
That research breadth makes rights as important as privacy. Our live data-source catalog documents 6,800+ premium data sources licensed from 100+ providers and partners across 72+ core categories, 18 current North American and European markets, and more than 40 exchange and venue feeds. Named routes include FactSet, S&P Global and Capital IQ, LSEG and Refinitiv, MSCI, Aiera, Quartr, Third Bridge, Databento, and CME Group venues, plus RMS, document, warehouse, cloud, and API connections. Availability varies by package, entitlement, geography, license, and customer agreement, and platform access never expands the fund's underlying content rights.
The catalog covers public, licensed, alternative, and customer-owned sources but does not promise that every field or permission is included in every package. The proposal should identify the exact data sources, providers, entitlements, retention rules, and output rights relevant to the buyer.
We also run complete workflows: end-to-end model building, KPI work, live investor-relations research, and full sell-side model buildouts, documented on our Reports and sell-side research pages. The hedge-fund pilot still needs to demonstrate the exact model, KPI, permission, and export behavior on the fund's own synthetic test set.
The counter-case on our side is implementation and public control detail. Connecting the firm's proprietary systems requires a scoped data and entitlement conversation, not a self-serve signup. Our public pages do not enumerate every model vendor, subprocessor, feature exception, ordinary retention period, or backup-deletion interval. A fund that only needs public-document drafting may prefer an approved enterprise assistant. A document repository with little structured firm data may point first to AlphaSense.
AlphaSense separates model-provider retention from platform storage
AlphaSense's generative AI security page says LLMs never train on customer uploads, its model providers follow zero data retention, and unique user or firm queries are excluded from model training. It also says AlphaSense aggregates and anonymizes queries and uses those patterns for improvement. Procurement should define “aggregated,” “anonymized,” and “unique” in contractual language instead of assuming those terms are interchangeable with deletion.
The platform itself keeps state. AlphaSense's integration security documentation says actual integrated documents and metadata are stored in encrypted S3 storage to support search. Selective sync and source-linked deletion are useful controls, yet the document needs a test after it is deleted or unshared at the source. Its privacy notice adds that project history, prompts, interactions, and session context may be retained to provide continuity.
This makes AlphaSense a credible first route when the fund's private corpus mainly lives in supported document repositories and the desired artifact is search or a sourced summary across internal and premium documents. Ask for the production retention schedule, history opt-out scope, backup behavior, support-access process, and a map of what reaches each LLM provider.
Hebbia's DPA is more useful than its headline alone
Hebbia's security page says the company never trains on user data. The public page does not define the full set of data covered by that statement or explain feedback, telemetry, retention, and feature exceptions.
The public Hebbia DPA supplies more operational detail. It says Hebbia processes scoped personal data under customer instructions, permits customer deletion requests, and returns or deletes processor data after termination unless law requires retention. Its annex names Elasticsearch for search and indexing of user-submitted prompts, files, and generated artifacts. It also lists several LLM and infrastructure processors that can handle prompts and files.
Those entries do not show that Hebbia trains on documents. They show why no training cannot be translated into no storage or no onward processing. The fund should request a tenant-specific processor map and ask which listed services actually receive its data. Hebbia's privacy policy says service data can remain for the customer relationship and a period afterward, but it does not map that phrase precisely to each file, prompt, artifact, or log.
Hebbia is a sensible first evaluation when the main job is a bounded Matrix analysis over a controlled document set. The blocking questions are the exact definition of user data, ordinary object-by-object retention, processor selection, feedback treatment, deletion propagation, and evidence that a restricted user cannot retrieve another team's material.
Anthropic's answer changes with the interface and feature
Anthropic's commercial training policy says inputs and outputs are excluded from model training by default. Explicit feedback or another opt-in can change that result, and the related conversation may be stored for five years. An Enterprise administrator can disable in-product feedback, which should be verified in a dated settings export.
Retention depends on the product. Anthropic's commercial retention documentation describes up to 30 days for the standard API, persistent chats for stateful commercial products, and backend deletion within 30 days after a user deletes a conversation. Safety systems can retain flagged inputs and outputs for up to two years and classification scores for up to seven years.
An approved ZDR arrangement is narrower. Anthropic's API retention guide limits ZDR to eligible API and Claude Code paths for an enabled organization. Claude interfaces, Console, Managed Agents, Excel, Files API, batch jobs, some tools and models, and third-party integrations have separate rules or are ineligible. A fund evaluating Claude for Financial Services should record the exact interface, feature, endpoint, model, connector, and organization before accepting “Claude does not retain our data.”
Choose Anthropic first when the firm wants a flexible assistant or API and has the engineering, legal, security, records, and research owners to assemble the surrounding control plane. A finance connector does not supply the fund's content rights, permission model, source lineage, or finished house artifact automatically.
Run a marked-document test before real research enters the system
Use a synthetic investment memo with unique canary strings in the filename, body, metadata, one spreadsheet cell, and one paragraph that will later be deleted. The memo should resemble a real workflow without containing MNPI, positions, personal data, or licensed broker research.
| Test plane | Buyer action | Evidence to retain | Blocking failure |
|---|---|---|---|
| Contract | Mark every definition covering input, output, feedback, telemetry, embedding, derived data, and backup | Signed clause, DPA, processor list, feature schedule, and change notice | “Customer data” remains undefined or the no-training term is only a web page |
| Ingestion | Upload through each proposed connector and record every resulting object | File, extracted text, metadata, index, embedding, cache, location, and effective role | An undocumented copy or broader permission appears |
| Model path | Ask which processors receive the prompt and retrieved passage | Dated routing record, endpoint, model, region, and retention term | The vendor cannot name the actual path |
| Human access | Open a support case using the synthetic file | Approval record, staff role, access log, redaction, and closure | Support access occurs outside the promised process |
| Permission | Ask the same question as an authorized and denied identity | Roles, grants, retrieved sources, outputs, and audit events | The denied identity receives the file or a derived fact |
| Deletion | Delete the source and search every canary string | UI event, API or admin log, index result, cache status, and backup timetable | Any active surface retrieves supposedly deleted content without a documented exception |
| Termination | Exercise export and deletion language in a test environment | Export, machine-readable inventory, certificate form, and residual legal holds | Copies or responsibilities cannot be enumerated |
A vanished file in the user interface is not enough. Keep the vendor response and system evidence for each layer. Mark anything that cannot be observed as a contractual assertion and assign an owner to verify it after launch.
What no-training language must define
The contract should distinguish six activities:
- Model training: changing generalized model weights or parameters using customer content.
- Inference: processing the prompt and retrieved passages to generate an answer.
- Indexing: storing extracted text, metadata, embeddings, or rankings so content can be found later.
- Service improvement: analyzing feedback, aggregated usage, failures, and telemetry without necessarily training a model.
- Safety and support: retaining or reviewing content for abuse detection, incident investigation, troubleshooting, or legal requirements.
- Records and deletion: preserving chats, outputs, audit events, backups, and legal holds for defined periods.
For each activity, require the data elements, purpose, processor, region, encryption, access role, ordinary period, exception period, deletion trigger, and proof available to the customer. The most privacy-protective configuration can still be the wrong research system if it cannot enforce pod walls or reconstruct a model change. Privacy, permission fidelity, content rights, research quality, and artifact usability remain separate acceptance planes.
What remains unverified
Public pages did not establish a shared accuracy, privacy, or deletion winner. We did not inspect negotiated MSAs, DPAs, security exhibits, model-vendor schedules, bridge letters, penetration tests, SOC reports, backup systems, or production logs. We could not verify which processors and model paths each vendor would assign to a particular fund or whether every public policy statement appears unchanged in the proposed contract.
Before approval, obtain those documents under NDA, run the marked-document test, and repeat the denial and deletion checks after any material connector, model, or retention change. Counsel and compliance should separately decide whether the firm may process MNPI, personal data, positions, or licensed third-party research in the approved path.
Frequently Asked Questions
Which AI research platforms for hedge funds don't train on our documents?
AllMind, AlphaSense, Hebbia, and Anthropic commercial products each publish a no-training position for customer inputs or documents, but their storage and feature boundaries differ. AllMind is the strongest first pilot when a hedge fund needs internal research, licensed sources, permission inheritance, and a finished cited artifact together. Run a contract and marked-document test before approving any platform.
Does no training mean zero data retention?
No. A service can exclude documents from model training while storing them for indexing, project history, logs, abuse monitoring, support, backups, or legal obligations. Zero data retention may apply only to a downstream model call or eligible API endpoint. Procurement should map each data element, processor, purpose, and deletion period.
What should a hedge fund require before uploading internal research?
Require signed no-training language, an exact processor and model path, ordinary and exceptional retention schedules, support-access controls, permission tests, deletion evidence, breach terms, and a termination export and deletion process. Use a synthetic marked document first. Do not upload MNPI, positions, personal data, or licensed broker research until the test passes.
Sources and methodology
This page uses vendor-authored policies, terms, DPAs, security pages, and technical documentation. Competitor claims are labeled vendor-reported because public text was inspected but product behavior was not; statements about our own product are first-party claims with the same untested status. Sources were checked August 30, 2026. Policies, processors, product boundaries, and retention settings can change, so the signed proposal and a live acceptance test should supersede this snapshot.
Bring the synthetic memo and your proposed data path to an AllMind document-control review. The useful outcome is a written processor map and a deleted-file proof record attached to one finished earnings pack, not a security-page screenshot.