What's new
  • The default language of any content posted is English.
    Do not create multi-accounts, you will be blocked! For more information about rules, limits, and more, visit the Help page.
    Found a dead link? Use the report button!

Spent two months collecting AI training data and hit quite a few proxy pitfalls

Kerry

Member
Messages
34
Reactions
0
I started collecting public web data for model pre-training at the end of last year. At first I went with some cheaper providers. Once concurrency went up, success rates dropped below 60%, and IPs were frequently flagged. After several rounds of switching, things gradually stabilized.

I tried pure residential, ISP, mobile, and mixed setups. Residential performed clearly better against Cloudflare, though traffic cost more. ISP worked better for tasks that needed session persistence. Mobile sometimes passed more easily in certain regions. I’m currently mainly using Helodata’s residential pool (helodata.com). City-level targeting is accurate and integration is straightforward, so no major failures lately.

Anyone else doing similar large-scale collection? What success rate are you able to maintain? Any particularly effective rotation or retry strategies? Would love to hear real numbers.
Post automatically merged:

Just to add, the one I’m currently using is this: [


] Success rate is a bit more stable than before, for reference.
 
Reacted by:
Top