Kerry
Member
- Messages
- 34
- Reactions
- 0
I started collecting public web data for model pre-training at the end of last year. At first I went with some cheaper providers. Once concurrency went up, success rates dropped below 60%, and IPs were frequently flagged. After several rounds of switching, things gradually stabilized.
I tried pure residential, ISP, mobile, and mixed setups. Residential performed clearly better against Cloudflare, though traffic cost more. ISP worked better for tasks that needed session persistence. Mobile sometimes passed more easily in certain regions. I’m currently mainly using Helodata’s residential pool (helodata.com). City-level targeting is accurate and integration is straightforward, so no major failures lately.
Anyone else doing similar large-scale collection? What success rate are you able to maintain? Any particularly effective rotation or retry strategies? Would love to hear real numbers.
Just to add, the one I’m currently using is this: [
] Success rate is a bit more stable than before, for reference.
I tried pure residential, ISP, mobile, and mixed setups. Residential performed clearly better against Cloudflare, though traffic cost more. ISP worked better for tasks that needed session persistence. Mobile sometimes passed more easily in certain regions. I’m currently mainly using Helodata’s residential pool (helodata.com). City-level targeting is accurate and integration is straightforward, so no major failures lately.
Anyone else doing similar large-scale collection? What success rate are you able to maintain? Any particularly effective rotation or retry strategies? Would love to hear real numbers.
Post automatically merged:
Just to add, the one I’m currently using is this: [
Helodata | Residential, ISP & Mobile Proxies for AI & Web ScrapingPremium residential, ISP, mobile & datacenter proxies for AI and web scraping. 80M+ ethically sourced IPs across 195 countries, from $2.50/GB.
|
] Success rate is a bit more stable than before, for reference.
Reacted by: