crawler

About this crawler

If you’ve found this page, you’ve probably seen EditfloBot in your server logs and want to know who it is and what it’s doing.

Who

I’m Van, a wedding photographer in Queensland. I build software for wedding photographers, and this crawler is part of that work. It’s a small operation — one person, not a company with a data division.

What it’s for

I’m building a tool that helps wedding photographers write blog posts about the weddings they shoot. One of the fiddliest parts of that job is crediting the other suppliers who worked the day — the florist, the celebrant, the venue, the hair and makeup artist. Photographers want to credit everyone properly, but hunting down eight businesses’ Instagram handles at 11pm is where good intentions go to die, and vendors end up uncredited.

So the tool needs a list: wedding suppliers, what they do, roughly where they work, and their website and Instagram handle. Type three letters, get the right business with the right handle attached. That’s what this crawler builds.

The same list underpins a wedding vendor directory I’m developing.

The short version: this crawler exists so that wedding suppliers get tagged and linked more often, not less.

What it collects

Public business information, the kind you’d put on a business card:

  • Business name
  • What category of wedding service you provide
  • The region you work in
  • Your website address
  • Your Instagram handle, where your site links to it
  • A business contact email or contact form address, where one is published

That’s the whole list.

What it doesn’t do

  • It doesn’t copy your content. No photographs, no portfolio images, no blog posts, no page text beyond the contact details above. Nothing you’ve written or shot is stored or republished.
  • It doesn’t log in to anything. Public pages only. No accounts, no paywalls, no members’ areas.
  • It doesn’t collect personal information. Business contact details only. If a page lists a staff member’s personal mobile or private address, it isn’t captured.
  • It doesn’t crawl Instagram. Handles are only ever read from links you’ve put on your own website.

How it behaves

  • One request per second, maximum, to any single site. Usually far less. It’s lighter traffic than one person browsing your site.
  • It reads your robots.txt first, every time, and does what it says.
  • It caches. Once a page has been read, it isn’t fetched again on later runs.
  • It identifies itself honestly, which is why you’re reading this.

The User-Agent string is:

EditfloBot/1.0 (+https://editflo.com/crawler)

How to keep it off your site

Add this to your robots.txt:

User-agent: EditfloBot
Disallow: /

It will be respected on the next run, and no further pages will be fetched.

Or just email me and I’ll exclude you — no explanation needed, no questions asked.

How to be removed from the list

Email me and I’ll delete your record. Same deal: no explanation needed.

If instead you’d like your listing corrected — wrong category, wrong region, outdated handle — email me that too. I’d rather have it right.

Contact

hello@editflo.com

I read these. A real person will reply.