Human data · Latin America
Pluri runs its own collection infrastructure across the region: speech, imagery and real-world task data, captured by identity-verified contributors, validated by automated QA and human review, and paid on local rails from Pix to SPEI. Every delivery ships with per-record provenance.
record
ACCEPTED
the provenance pipeline, live · illustrative values
The gap
Published research on Portuguese language models shows that the gain from native-language pretraining comes from cultural and domain knowledge, not grammar. The same logic holds across the region. What models are missing is not syntax. It is Latin American context: how people speak, pay, shop, complain and fill out forms.
That context does not exist in scraped web data. It has to be collected from people: with regional accent balance, real devices, real environments and documented consent.
Pluri exists to supply it, at contract grade.
Capabilities
Custom collection against your spec. Fresh, task-generated data from real contributors. Nothing scraped, nothing synthetic.
Scripted prompts, spontaneous conversation, wake words, code-switching, noisy environments. Accent quotas balanced across the region, in Brazilian Portuguese and Latin American Spanish.
Real-world photos, receipts, screens, shelves and street scenes. Captured on the contributor's own device, validated at submission time.
Guided real-world tasks on Latin America's digital rails: instant payments, delivery, e-commerce, government services. Step-level capture, with consent.
Industries
Datasets and standing panels are scoped per industry. What we capture is what the market actually shows people, and what people actually do about it.
Fees and rates as displayed, credit offers received, onboarding and payment flows across banks and fintechs.
Order receipts, baskets and prices paid, promos seen, delivery times as experienced across delivery apps.
Search results, product pages, carts and checkouts as shown to real shoppers, across competing marketplaces.
Shelf and street photography, in-store price checks and paper receipts from physical retail, digitized and validated.
Fare quotes, surge behavior and wait times across ride-hailing and mobility apps, captured on both sides of the trip.
Plan and price screens, paywalls, churn and win-back offers across streaming and subscription products.
Provenance
This is what one delivered unit looks like. Not a summary metric over the dataset: the audit trail of a single record, the way it ships.
record plr_a91c04…7e4
SPEECH · PT-BR
ACCEPTED
Ask any vendor for this view of one record. That question is the entire audit.
Field values above are illustrative. The structure is contractual.
Pipeline
The same four-stage pipeline runs every modality. Change the task, keep the controls.
Recruiting from an owned, identified contributor base, filtered by country, city, demographic and accent quota. No marketplaces, no scraping, no anonymous crowd.
Contributors complete the task in a conversational collector that validates at submission time: wrong city, wrong format or missing fields never enter the pipeline.
Automated checks for authenticity, spec match and cross-contributor duplication, then a human review gate on every batch. Uncertain units go to quarantine, not to you.
Contributors are paid on local instant rails, Pix and country equivalents, after client acceptance, with tax reporting handled. Paid contributors return. Returning contributors are the quality flywheel.
Continuous panels
The same machine, left running. Consented cohorts of 500 to 2,000 people per industry submit screen-level evidence from their own devices every month: prices paid, offers seen, flows completed. Longitudinal, verified, redacted.
A monthly delivery combining the record-level dataset and an aggregated read: price and promo movements, share shifts, flow changes. Same cohort over time, so a movement is a movement, not sampling noise.
Syndicated by default: every subscriber receives the same deliverable, and no one can commission collection aimed at a specific company. A separate license makes the underlying screens available for model training.
Proven in production: our standing screen-collection program runs at a contracted volume of 20,000 validated units per month.
For data platforms
If you sell data to AI labs and need Latin American supply that survives your QA bar, this is the product.
Your spec, our field. Send collection requirements and get a costed pilot proposal within 72 hours. Daily batch manifests, verdicts returned on your template, rejected units replaced.
Contract a monthly throughput floor by modality, country, city and demographic, with burst on demand. Your roadmap gets guaranteed Latin American supply instead of spot-market risk.
Our field operation has delivered to global data platforms and research companies, under NDA, since 2019.
Compliance
Consent logged with version, timestamp and scope. Purpose-bound collection, opt-out honored and recorded, retention and deletion per contract. Brazil's LGPD as the baseline, aligned with privacy regimes across the region. Data processing agreement as standard.
Personal identifiers are masked at ingestion, before storage and before delivery, with a per-record redaction log. Uncertain redactions are quarantined for human review, never shipped.
A registered company in the region, enforceable contracts, and contributor support in Portuguese and Spanish. When something needs resolving, there is someone to call. In this market, that is rarer than it should be.
Contributors are identified and paid per accepted unit on local rails, with tax reporting handled. Payment terms are stated before the task starts. Disputes have a channel and an SLA.
Start
One page is enough: modality, volume, demographic targets and your QA bar. We reply with scope, timeline and unit pricing.