The problem is not building the model. It is reaching the data: it lives in PDFs, in spreadsheets, or in systems nobody wants to touch.
When a Brazilian healthtech pitches a health plan or a hospital, the expectation is that the main challenge is the algorithm: predictive models, AI, accuracy. The reality is different. Even sophisticated health plans have a significant share of admissions without structured data. Reports live in open PDFs. Historical clinical decisions are locked inside closed medical records systems.
A phrase I heard in more than one conversation, from different sources: "the data is in PDFs at the health plans." It is not a lack of data. It is a lack of usable structure.
Three practical implications for anyone selling data solutions to healthcare in Brazil:
- Whoever solves the structured extraction bottleneck wins before whoever optimizes the model does. OCR, NER on medical reports, parsing badly formatted codes: the tedious work is where the value is.
- The startup that couples into the existing platform wins over the one that asks for a replacement. Health plans do not want to swap systems, they want to add capacity.
- The health plan buys savings with little work, not technology with a lot of work. If the pitch requires a six month IT project to get started, the pitch has already lost.
This holds true beyond healthcare, too: the pattern repeats in any sector with consolidated legacy systems (industry, agribusiness, government, large enterprises). The bottleneck is always operational before it is technical.