#Eroforze Lead Scraper: Operator Manual

This manual covers everything needed to install, run, configure, monitor and troubleshoot the lead scraper. You don't need to read the code.


#1. What the scraper does

The scraper is a background service. It runs in cycles, by default one every 60 minutes. Each cycle has two phases.

#Phase 1: Discovery (new companies)

  1. It searches the web for each query in config.json, for example "software development company in Chennai".
  2. It also adds any seed domains you list yourself.
  3. It removes directories, job boards, social networks and marketplaces (LinkedIn, Justdial, Naukri, Clutch, Glassdoor and so on).
  4. It asks the CRM whether each remaining domain already exists. Existing companies are skipped.
  5. It visits each new website: the homepage plus up to 5 pages such as Contact, About, Team and Careers.
  6. If the site has a company name and at least one email or phone number, it uploads it to the CRM as a new lead, tagged with the query that found it.

#Phase 2: Enrichment and verification (existing leads)

  1. It reads every lead in the CRM through the API, oldest first.
  2. It picks leads that are missing something and haven't been checked in the last 30 days.
  3. For each one it:
    • finds the website if the lead only has a company name,
    • crawls the website for emails, phones, address, country, company LinkedIn and careers page,
    • finds the person's LinkedIn if the lead is a person,
    • finds HR or recruiting people at the company (team page first, then public LinkedIn search results),
    • verifies the email address,
    • checks that the website is still live.
  4. It sends only the missing fields back to the CRM. Nothing already in the CRM is ever overwritten, except the verification fields the scraper owns (email status, "verified", website status).
  5. Each HR person found becomes a separate lead linked to the company, and appears as an HR contact on the Companies page.

#Where data goes

Every result is written both to the CRM and to files on the server (data/). If the CRM is unreachable, uploads wait in a local outbox and are resent automatically on the next cycle, so nothing is lost.

#What it does not do

  • It does not use AI, paid data providers or paid APIs.
  • It never logs in to LinkedIn or scrapes LinkedIn directly. It only reads public search-engine results.
  • It never sends emails. Even the optional mailbox check stops before a message is sent.

#2. Before you start

#Server requirements

ItemRequirement
Operating systemAny 64-bit Linux (Ubuntu 22.04/24.04 recommended). macOS works for testing.
CPU / RAM1 vCPU and 1 GB RAM minimum; 2 vCPU / 2 GB recommended.
Disk2 GB free (the Docker images plus result files).
SoftwareDocker with Docker Compose v2 (recommended), or Node.js 22.18 or newer.
Network outPort 443 (CRM API and websites) and 80 (some websites). Port 25 only if you enable mailbox checks.
Network inNone. The scraper opens no ports.

Install Docker on Ubuntu if it isn't there:

curl -fsSL https://get.docker.com | sudo sh
sudo usermod -aG docker $USER   # then log out and back in
docker compose version          # should print v2.x

#What you need from the CRM

An API key with read and write access to leads. Section 3.1 explains how to create one.


#3. Installation

#3.1 Create the API key in the CRM

  1. Sign in to the CRM and open Settings → API Keys.
  2. In Generate API Key, fill in:
    • Key name: something you'll recognise, for example Lead scraper (server 1).
    • Source label: lead-scraper. Use the same value as SCRAPER_SOURCE in .env (the default is lead-scraper).
    • Product: the product new leads should belong to. Leave it as Default if unsure.
    • Access: choose Scraper ("Send leads and check for duplicates").
  3. Click the generate button. The CRM shows "Copy your new key now". Copy the key (it starts with efz_). It's shown only once. If you lose it, revoke it and create a new one.

One key per server is best practice. You can see each key's activity and revoke it on its own.

#3.2 Copy the folder to the server

From the machine that has the scraper folder:

rsync -av --exclude node_modules --exclude data scraper/ user@your-server:~/scraper/
# or:  scp -r scraper user@your-server:~/

#3.3 Run the installer

ssh user@your-server
cd ~/scraper
./setup.sh

The installer:

  1. creates .env from .env.example if it doesn't exist,
  2. asks you to paste the API key and saves it in .env,
  3. generates a random secret for the private search engine (SearXNG),
  4. builds the scraper image and starts two containers that restart automatically, including after a server reboot:
    • eroforze-scraper: the scraper itself,
    • eroforze-searxng: a private search engine only the scraper can reach,
  5. runs a connection check and prints the result.

A healthy check looks like this:

  OK   CRM API: https://site-crm-data.eroforze.com accepts key "Lead scraper (server 1)"
  OK   SearXNG: http://searxng:8080 answers JSON searches
  OK   Web search: 10 results
  OK   DNS (MX lookups): 5 mail servers for gmail.com
  OK   Outbound port 25: blocked (fine, SMTP_VERIFY is off)

If anything shows FAIL, see Troubleshooting. The scraper keeps running either way.

#3.4 Installing without Docker

If Docker isn't available, setup.sh installs with Node.js instead:

  • Run as root: it installs a systemd service called eroforze-scraper that starts on boot.
  • Run as a normal user: it starts the scraper in the background and writes logs to data/scraper.log.

Without Docker there's no private search engine, so search falls back to public DuckDuckGo. That allows only a handful of searches per hour, which limits discovery, HR lookups and LinkedIn lookups. Use Docker for production.


#4. Your first run (safe test)

Before letting the scraper write to the CRM, do one dry run. It crawls and saves everything locally but sends nothing to the CRM.

cd ~/scraper
docker compose run --rm -e RUN_ONCE=true -e DRY_RUN=true scraper

When it finishes, inspect what it would have uploaded:

ls data/results/
less data/results/$(date +%F).jsonl

Each line is one record:

  • "type":"new-company": a company discovery would add,
  • "type":"enrich": the fields it would fill on an existing lead (filled, patch),
  • "type":"hr-contact": an HR person lead it would create,
  • "type":"cycle": the summary of the run.

When you're happy, start it for real:

docker compose up -d
docker compose logs -f scraper

Tip: for the first live run, set "maxNewCompaniesPerCycle": 10 and "maxLeadsPerCycle": 25 in config.json so you can review a small batch in the CRM first.


#5. Day-to-day commands

Run all commands from the scraper folder.

#With Docker (normal setup)

TaskCommand
Watch live logsdocker compose logs -f scraper
Last 200 log linesdocker compose logs --tail 200 scraper
Is it running?docker compose ps
Stop everythingdocker compose down
Start againdocker compose up -d
Apply changes to config.json or .envdocker compose up -d --force-recreate scraper
Run one cycle now and exitdocker compose run --rm -e RUN_ONCE=true scraper
Dry run (nothing sent to the CRM)docker compose run --rm -e RUN_ONCE=true -e DRY_RUN=true scraper
Connection checkdocker compose run --rm scraper node src/check.ts
Search-engine logsdocker compose logs --tail 100 searxng

Stopping is graceful. On docker compose down or restart, the scraper finishes the lead it's working on and then exits. Anything not yet sent is either already in the CRM or in the outbox.

#Without Docker (systemd)

TaskCommand
Live logsjournalctl -u eroforze-scraper -f
Statussystemctl status eroforze-scraper
Stop / start / restartsudo systemctl stop|start|restart eroforze-scraper
One cycle nowRUN_ONCE=true node src/index.ts
Dry runnpm run dry-run
Connection checknpm run check

#6. Choosing what it finds (config.json)

config.json controls what the scraper looks for and how much it does per cycle. After editing, apply the change with docker compose up -d --force-recreate scraper.

The file must stay valid JSON: double quotes, no trailing commas, no comments. If unsure, paste it into any online "JSON validator" first.

#6.1 Discovery queries

"discovery": {
  "enabled": true,
  "maxNewCompaniesPerCycle": 50,
  "revisitDays": 90,
  "queries": [
    {
      "query": "software development company",
      "locations": ["Bangalore", "Chennai", "Hyderabad"],
      "country": "IN",
      "industry": "Software Development",
      "tags": ["it-services"],
      "product": "edgeryt-hire",
      "pages": 2
    }
  ],
  "seeds": [],
  "blocklist": []
}
FieldWhat it does
queryThe search phrase. Describe the type of company, for example "logistics company", "manufacturing company hiring" or "fintech startup".
locationsEach location is searched separately as "query in location". Leave it out or use [] to search the query alone.
countryISO code (IN, AE, US, GB, SG…) used only if the website doesn't reveal its country.
industrySaved in the lead's Industry field.
tagsAdded to customFields.tags together with discovered. Use them to filter or build campaigns later.
productCRM product key for these leads. If omitted, SCRAPER_PRODUCT from .env is used, then the API key's product.
pagesHow many pages of search results to read per location (about 10 results per page). Start with 1–2.

Other discovery settings:

SettingMeaning
maxNewCompaniesPerCycleUpper limit of new companies added per cycle.
revisitDaysA domain discovery has already looked at is ignored for this many days.
enabledfalse turns discovery off; only enrichment runs.

How queries are worked through. Each cycle goes through the queries in order until it has enough candidates. Search results are cached for 14 days and checked domains are remembered for 90 days, so later cycles naturally move on to queries and locations not yet covered. Long query lists are fine; they're simply spread over several cycles.

Writing good queries:

  • Be specific about the company type: "SaaS company" beats "companies".
  • Add intent words that suit your product: "hiring", "startup", "careers".
  • Use cities rather than whole countries for better local results.
  • Avoid queries that mainly return lists or articles, such as "top 10 IT companies". These pages are filtered out, so they waste searches.

#6.2 Seed domains

Seeds are companies you already know you want added and enriched, such as event attendees or a purchased list of websites:

"seeds": [
  { "domain": "acme.com", "country": "IN", "industry": "Manufacturing", "tags": ["expo-2026"] },
  { "domain": "https://www.example.ae", "country": "AE", "tags": ["partner-referral"] }
]

Seeds are processed before search queries. They go through the same checks: skipped if already in the CRM, and added only if the site shows a name and a way to contact the company.

#6.3 Blocklist

The scraper already ignores well-known directories, job boards, social networks, news sites and marketplaces. Add anything else that keeps appearing but isn't a real prospect:

"blocklist": ["competitor.com", "somedirectory.in", "examplelistings"]

An entry with a dot blocks that exact domain. An entry without a dot blocks every domain containing that name ("examplelistings" blocks examplelistings.com, examplelistings.co.in and so on).

#6.4 Enrichment settings

"enrichment": {
  "enabled": true,
  "maxLeadsPerCycle": 200,
  "refreshDays": 30,
  "product": "",
  "source": "",
  "fillCompanyEmail": true,
  "findHrContacts": true,
  "createHrLeads": true,
  "maxHrContactsPerCompany": 2,
  "findPersonLinkedin": true,
  "findMissingWebsites": true
}
SettingMeaning
maxLeadsPerCycleHow many leads are enriched per cycle.
refreshDaysAfter a lead is checked, it's left alone for this many days, even if nothing was found. Emails are also re-verified after this period.
productOnly enrich leads of this product (key or name). Empty means all products.
sourceOnly enrich leads from this source, for example csv-import. Empty means all sources.
fillCompanyEmailFor company leads (no person name) with no email, use the best company inbox as the email: HR inbox first, then a general inbox such as info@.
findHrContactsLook for HR and recruiting people at each company.
createHrLeadsCreate a separate lead for each HR person found. If false, only the company lead's hr_name / hr_title fields are filled.
maxHrContactsPerCompanyMaximum HR people added per company.
findPersonLinkedinFor person leads without a LinkedIn URL, search for their profile.
findMissingWebsitesFor leads with a company name but no website, search for the official website.

Leads with status Archived or Lost are never enriched.

#6.5 Crawl settings

"crawl": {
  "maxPagesPerSite": 6,
  "maxConcurrency": 10,
  "maxRequestsPerMinute": 120,
  "respectRobotsTxt": true,
  "timeoutSecs": 30
}
SettingMeaning
maxPagesPerSitePages read per website, homepage included. Only contact, about, team, careers, leadership and imprint-style pages are followed.
maxConcurrencyWebsites fetched in parallel.
maxRequestsPerMinuteOverall page-fetch speed limit. Each site is also limited to one page per second.
respectRobotsTxtObeys each site's robots.txt. Keep this true.
timeoutSecsA page that takes longer than this is skipped.

#7. Settings reference (.env)

.env holds the API key and switches. After editing, apply with docker compose up -d --force-recreate scraper.

SettingDefaultMeaning
EROFORZE_API_KEY(none)Required. The CRM API key (efz_…).
EROFORZE_API_URLhttps://site-crm-data.eroforze.comCRM API address.
SCRAPER_SOURCElead-scraperSource label on every lead the scraper creates. Shown in Settings → Data Sources and usable as a filter.
SCRAPER_PRODUCT(empty)Product key for new leads. Empty means the API key's product, then the workspace default.
CYCLE_INTERVAL_MINUTES60Pause between the end of one cycle and the start of the next.
RUN_ONCEfalsetrue runs one cycle and exits (for cron or testing).
DRY_RUNfalsetrue saves locally but sends nothing to the CRM.
SEARCH_ENABLEDtruefalse disables all web searches (no discovery by search, no HR or LinkedIn lookups, no missing-website lookups). Seeds and website crawling still work.
SEARCH_ENGINESsearxng,duckduckgoTried in order; falls back to the next one if the first fails or rate limits.
SEARXNG_URLset by DockerAddress of the private search engine. Leave empty in .env when using Docker.
SEARCH_DELAY_MS8000Minimum gap between searches on public DuckDuckGo. SearXNG uses a third of this. Raise it if searches get blocked.
SEARCH_CACHE_DAYS14How long search results are reused.
SEARXNG_SECRETgeneratedSecret for the private search engine. Don't share it or change it while running.
SMTP_VERIFYfalsetrue turns on mailbox-level email checks and HR email guessing. See section 11.
SMTP_HELO_DOMAINlocalhostDomain the scraper introduces itself with to mail servers. Use a real domain you own.
SMTP_FROMverify@localhostSender address used in the check (no email is ever sent). Use an address on that domain.
SMTP_TIMEOUT_MS12000How long to wait for a mail server.
PROXY_URLS(empty)Comma-separated proxies for website crawling, for example http://user:pass@1.2.3.4:8000.
USER_AGENTa desktop ChromeBrowser identity sent to websites.
LOG_LEVELinfodebug shows every search, crawl failure and SMTP reply.
DATA_DIRdataWhere local files are kept.

#8. Seeing the results in the CRM

#8.1 Data Sources

Settings → Data Sources shows a row for the scraper's source label (lead-scraper). It lists the number of requests, leads received, New, Updated and the last upload time. Use it to confirm the scraper is alive.

#8.2 Finding scraper leads

  • Leads page, filter Source = lead-scraper: every company and HR person the scraper created.
  • Leads that already existed keep their original source. The scraper only fills their gaps, and the changes appear in each lead's audit log under the API key's name.

#8.3 Fields the scraper fills

Standard fields (filled only if empty): Email, Phone, Website, Company, Country, City, State, LinkedIn URL. Email status and Verified at are set by verification.

Custom fields on company leads and existing leads:

Custom fieldExampleMeaning
hr_emailcareers@acme.comHR or recruitment inbox found on the website.
hr_name, hr_titlePriya Sharma, Senior HR Business PartnerHR person found for the company.
company_emailinfo@acme.comGeneral inbox.
company_linkedinhttps://www.linkedin.com/company/acmeCompany LinkedIn page.
careers_urlhttps://acme.com/careersCareers or jobs page.
tagsdiscovered, it-servicesTags from the discovery query.
lead_typecompany / personWhat kind of lead the scraper created.
discovered_viasearch: software development company in ChennaiThe query or seed that found it.
discovered_atISO dateWhen it was discovered.
search_locationChennaiThe search location that found it.
country_sourcewebsite address, domain, phone code, page language, search locationWhere the country came from.
country_detectedAEThe website's address says a different country from the one in the CRM. Worth a manual look.
verifiedtrue / falseEmail confirmed good or bad. Shown as Email verified on the Companies page.
email_checklisted on the company website, mail server foundWhy the email got its status.
email_found_byname pattern + mail server checkThe email was guessed from the person's name and confirmed by their mail server.
website_statuslive / unreachableWhether the website loaded.
enriched_atISO dateWhen the scraper last checked this lead.
enrich_resultfound phone, country; still missing hr_emailWhat it found and what's still missing.

HR person leads created by the scraper have Full name, Title, LinkedIn URL, Company, Website and Country, plus:

Custom fieldMeaning
linkedin_rolehr or recruiting.
linkedin_profile_of"Name - Title" (used by the Companies page).
hr_linkedin / recruiting_linkedinTheir LinkedIn profile.
parent_leadLead ID of the company lead they were found for.
tagshr-contact, scraped.
found_viacompany website or linkedin search.

#8.4 Companies page

Open a company on the Companies page to see its contacts. HR people the scraper created appear at the top as named HR contacts with their LinkedIn, together with the HR inbox (hr_email) and the General inbox.

#8.5 Building campaigns from scraper leads

Use the audience filters (Source = lead-scraper, Industry, Country, Product). Only leads with a usable email and the right consent status are emailed, so check consent rules before sending to cold leads.


#9. Reading the logs

Example of a normal cycle:

INFO  Eroforze lead scraper starting {"api":"https://site-crm-data.eroforze.com","source":"lead-scraper","search":["searxng","duckduckgo"],"smtpVerify":false,"dataDir":"/app/data"}
INFO  Connected to the CRM as API key "Lead scraper (server 1)" (default product edgeryt-hire)
INFO  Search "software development company in Chennai": 8 website(s)
INFO  Crawling 8 new website(s)
INFO  Uploaded 5 lead(s): 5 new, 0 updated, 0 unchanged, 0 rejected
INFO  New company Creatah Software Technologies (creatah.com) career@creatah.com IN
INFO  Discovery done {"searched":1,"candidates":8,"alreadyInCrm":0,"notACompany":3,"created":5}
INFO  Crawling 20 website(s)
INFO  LD-7K3QX9MZ Freshworks: found phone, city, state, country, careers_url, company_linkedin, email, hr_name, hr_title; still missing hr_email
INFO  LD-2HF8PL0Q Dead Co: nothing new found; still missing phone, country, linkedin, hr_email, hr_contact
INFO  Uploaded 3 lead(s): 3 new, 0 updated, 0 unchanged, 0 rejected
INFO  Enrichment done {"checked":20,"updated":12,"fieldsFilled":64,"verified":15,"hrLeads":3}
INFO  Cycle finished {"discovery":{...},"enrichment":{...},"searches":42,"websitesCrawled":28,"minutes":6.4}
INFO  Next cycle in 60 minute(s)

What the summary numbers mean:

NumberMeaning
searchedDiscovery searches run.
candidatesNew domains worth crawling (not already in the CRM).
alreadyInCrmDomains skipped because the company is already a lead.
notACompanyWebsites skipped: unreachable, or no name or contact details.
createdNew company leads added.
checkedExisting leads looked at.
updatedLeads that gained at least one field.
fieldsFilledTotal fields filled.
verifiedEmails that got a verification status.
hrLeadsHR person leads created.
searchesWeb searches actually sent this cycle (cached ones aren't counted).

Warnings you may see, and what they mean:

MessageMeaningAction
duckduckgo is rate limiting us; pausing it for 30 minutesThe public search engine is blocking us for now.Normal now and then. If constant, make sure SearXNG is running (Docker setup).
searxng is rate limiting us…Every engine inside SearXNG is blocked.Raise SEARCH_DELAY_MS, reduce queries or pages.
… failed after 6 attempts … Saved … to the outboxThe CRM couldn't be reached after retries.Nothing to do; it's resent next cycle. Check the CRM if it repeats.
Update of LD-… refused (400): …The CRM rejected a change (for example an invalid value).Read the message; usually a single bad value.
Lead rejected: …One new lead failed validation (for example a malformed email).Usually harmless; the rest of the batch was saved.
API key rejected (401) / revokedThe key is wrong or was revoked.Create a new key and update .env (section 12.3).
missing leads:read, leads:writeThe key lacks permissions.Create a key with the Scraper or Full access preset.

#10. Files saved on the server

Everything lives in scraper/data/:

FileContentsSafe to delete?
results/YYYY-MM-DD.jsonlEvery new company, HR contact and enrichment (what was found, what was sent, whether the CRM accepted it), plus one summary line per cycle.Yes. It's your local history and backup. Keep or archive as you like.
outbox.jsonlUploads waiting for the CRM. Normally empty.No. Deleting it loses unsent data.
seen-domains.jsonDomains discovery has already checked, with dates.Yes, if you want discovery to re-check every domain.
search-cache.jsonRecent search results.Yes. Searches will simply run again.
scraper.log, scraper.pidOnly in no-Docker mode.Yes, when stopped.

Useful one-liners:

# today's new companies
grep '"type":"new-company"' data/results/$(date +%F).jsonl | wc -l

# everything sent about one lead
grep 'LD-7K3QX9MZ' data/results/*.jsonl

# cycle summaries for today
grep '"type":"cycle"' data/results/$(date +%F).jsonl

# export today's new companies as CSV (needs jq)
grep '"type":"new-company"' data/results/$(date +%F).jsonl \
  | jq -r '.lead | [.companyName, .website, .email, .phone, .country, .industry] | @csv'

#11. Email verification explained

Every email the scraper adds, and every existing email not verified in the last refreshDays, gets a status in the CRM's Email status field.

StatusMeaningSafe to email?
validThe domain receives mail and either the company publishes the address on its own website, or the mail server confirmed the mailbox (SMTP mode).Yes.
catch_allThe company's mail server accepts any address, so this one can't be confirmed (SMTP mode only).Probably; bounces are possible.
riskyDisposable email domain, or the mail server temporarily refused the check.Avoid or check by hand.
invalidBad format, the domain doesn't exist or has no mail server, or the mail server said the mailbox doesn't exist.No.
unknownThe domain receives mail, but the specific mailbox couldn't be confirmed.Use with care.

Two rules protect existing data:

  • An unknown result never replaces a status someone already set.
  • customFields.verified is set to true only for valid and to false only for invalid.

#Turning on mailbox checks (SMTP_VERIFY=true)

Turning this on adds three things:

  • a check with the company's mail server that the mailbox really exists,
  • detection of catch-all domains,
  • guessing personal and HR emails from a person's name (for example priya.sharma@, priya@, psharma@), accepted only when the mail server confirms the address and the domain isn't catch-all.

Requirements:

  1. Outbound port 25 must be open. Run the connection check; it reports "Outbound port 25: open/blocked". AWS, Google Cloud, Azure and many VPS providers block it by default. Some unblock it on request.
  2. Set SMTP_HELO_DOMAIN and SMTP_FROM to a real domain you control, ideally one whose server IP has a matching reverse-DNS (PTR) record. Without that, some mail servers refuse to answer.
  3. Don't use your main email-sending server's IP. Heavy checking can affect that IP's reputation.

#12. Maintenance

#12.1 Changing settings

  1. Edit config.json or .env.
  2. Run docker compose up -d --force-recreate scraper.
  3. Watch docker compose logs -f scraper for the "starting" line with your new settings.

#12.2 Updating the scraper code

Copy the new src/ (and package.json / package-lock.json if they changed), then rebuild:

cd ~/scraper
docker compose up -d --build

Your .env, config.json and data/ are kept.

To update the search engine image now and then:

docker compose pull searxng && docker compose up -d searxng

#12.3 Rotating or replacing the API key

  1. In the CRM, Settings → API Keys, create a new key (preset Scraper).
  2. Put it in .env as EROFORZE_API_KEY=….
  3. Run docker compose up -d --force-recreate scraper and confirm "Connected to the CRM as API key …" in the logs.
  4. Revoke the old key in the CRM.

#12.4 Making the scraper re-check specific leads

A lead is skipped for refreshDays after its last check. To re-check it sooner, clear its enriched_at custom field, either in the CRM or through the API:

curl -X PATCH "https://site-crm-data.eroforze.com/api/v1/leads/LD-7K3QX9MZ" \
  -H "Authorization: Bearer $EROFORZE_API_KEY" -H "Content-Type: application/json" \
  -d '{"customFields":{"enriched_at":null}}'

To re-check everything sooner, lower refreshDays temporarily (for example to 1), then set it back.

#12.5 Backups and moving to another server

  • Back up scraper/.env, scraper/config.json and, optionally, scraper/data/.
  • To move: copy the whole folder (without node_modules), run ./setup.sh on the new server, and stop the old one with docker compose down. Never run two copies with the same configuration at once. They'd do the same work twice. The CRM's duplicate protection prevents duplicate leads, but you'd waste searches.

#12.6 Running more than one scraper

To cover different markets in parallel, run separate copies of the folder on the same or different servers. Give each copy its own config.json queries, its own API key and, if you want them told apart in Data Sources, its own SCRAPER_SOURCE. Set enrichment.enabled to false on all but one copy, so only one enriches existing leads.

#12.7 Disk usage

Result files are small, typically a few MB per month. Container logs are capped at 50 MB for the scraper and 30 MB for the search engine. To remove old result files:

find data/results -name '*.jsonl' -mtime +180 -delete

#13. Troubleshooting

#The connection check fails

CheckLikely causeFix
CRM API FAIL 401Wrong or mistyped key.Re-copy the key into .env, recreate the container.
CRM API FAIL revokedThe key was revoked.Create a new key (section 12.3).
CRM API FAIL missing leads:read…The key was created with the wrong access.New key with the Scraper preset.
CRM API FAIL network errorThe server can't reach the CRM.Test curl https://site-crm-data.eroforze.com/api/v1/health; check firewall and DNS.
SearXNG FAIL unreachableThe search container isn't running.docker compose ps, then docker compose logs searxng, then docker compose up -d searxng.
SearXNG FAIL answered 403JSON output disabled.Make sure searxng/settings.yml lists json under search.formats, then docker compose restart searxng.
Web search FAIL no resultsAll search engines are blocking the server's IP right now.Wait 30–60 minutes; raise SEARCH_DELAY_MS; consider a server with a different IP. Crawling seeds and existing websites still works.
DNS FAILThe server can't resolve domains.Check /etc/resolv.conf or Docker DNS.
Port 25 FAILPort 25 blocked while SMTP_VERIFY=true.Set SMTP_VERIFY=false, or ask your provider to unblock port 25.

#The scraper runs but…

SymptomCause and fix
No new companiesCheck the logs for alreadyInCrm (the companies are already in the CRM) and notACompany (the sites have no contact details). Add more locations or queries, raise pages, or check searches aren't blocked. Domains are remembered for 90 days; delete data/seen-domains.json to re-check them.
"nothing new found" on most leadsThe leads' websites don't publish the missing data, or the sites are built entirely in JavaScript (the crawler reads plain HTML). Normal for a portion of leads.
Few HR contactsHR lookups depend on search. Make sure SearXNG works (connection check). Small companies often have no public HR profiles.
Few personal emailsWithout SMTP_VERIFY=true, personal emails are only added when published on the company's site. See section 11.
Wrong country on some leadsWhen a website has no address, the country comes from the domain (.in → India), the phone code, or the query's country. Look at country_source on the lead. Existing countries are never changed; disagreements are flagged in country_detected.
Wrong company nameIt's taken from the website's structured data, its site name or the page title. Fix it in the CRM; the scraper never overwrites it.
Leads in the wrong productSet product on the query, or SCRAPER_PRODUCT, or the API key's product.
Container keeps restartingdocker compose logs --tail 50 scraper. Usually a missing or invalid API key, or invalid JSON in config.json.
outbox.jsonl keeps growingThe CRM has been unreachable or failing for a while. Check the CRM and the API URL. Items are resent automatically once it's back.
High CPU or memoryLower crawl.maxConcurrency (for example to 5) and maxRequestsPerMinute.

#Getting more detail

Set LOG_LEVEL=debug in .env and recreate the container. You'll see every search query and result count, every page that failed to load, the reasons engines were skipped, and SMTP replies. Set it back to info afterwards; debug output is large.


#14. Safety, privacy and stopping in an emergency

#Stop immediately

docker compose down

To cut off CRM access completely, even if someone restarts the containers, revoke the API key in Settings → API Keys.

#Undoing a bad batch

  • Every lead the scraper created has the scraper's source label, so you can find and bulk-review them on the Leads page (filter by Source).
  • Every change to an existing lead is recorded in that lead's audit log, with the old and new values, under the API key's name.
  • data/results/*.jsonl lists exactly what was sent and when.

#Good practice

  • Consent and law: scraped contacts are cold leads. Follow the rules that apply to you (India's DPDP Act, GDPR for EU contacts, CAN-SPAM for the US, UAE PDPL). Always include an unsubscribe option; the CRM's campaigns already do.
  • Be polite to websites: keep respectRobotsTxt: true. Don't raise speed limits much beyond the defaults.
  • Search engines: automated searching can break search engines' terms. The private SearXNG plus caching keeps volume low, but large query lists mean more searches.
  • LinkedIn: the scraper never logs in to or crawls LinkedIn. Profile links come only from public search results and companies' own websites, and are accepted only when the result names both the person and the company.
  • Secrets: .env contains the API key. Don't commit it to git or share it. The folder's .gitignore already excludes it.

#15. Speed and capacity

These are rough figures for the default settings on a 2 vCPU server. Actual numbers depend on how fast websites respond and how often search engines limit you.

ActivityTypical speed
Website crawlingAbout 20 websites in 20–60 seconds (6 pages each, 10 in parallel).
Search (SearXNG)One search every 3–4 seconds, about 1,000 per hour at most. Cached results are free.
Search (DuckDuckGo only, no Docker)A few searches before a 30-minute pause.
Enrichment1–3 searches per lead (website, LinkedIn, HR) plus a crawl. Often 100–300 leads per hour.
DiscoveryEach query and location costs pages searches, then a crawl of the new domains.

Tuning tips:

  • More leads per day: raise maxLeadsPerCycle and maxNewCompaniesPerCycle, or lower CYCLE_INTERVAL_MINUTES. Watch the logs for rate-limit warnings.
  • Fewer search blocks: raise SEARCH_DELAY_MS, lower pages, or set findPersonLinkedin to false (it's the most search-hungry option).
  • Lighter on the server: lower maxConcurrency and maxRequestsPerMinute.

#16. Quick reference card

INSTALL          ./setup.sh
LOGS             docker compose logs -f scraper
STATUS           docker compose ps
CHECK            docker compose run --rm scraper node src/check.ts
STOP / START     docker compose down   |   docker compose up -d
APPLY CHANGES    docker compose up -d --force-recreate scraper
RUN ONE CYCLE    docker compose run --rm -e RUN_ONCE=true scraper
DRY RUN          docker compose run --rm -e RUN_ONCE=true -e DRY_RUN=true scraper
UPDATE CODE      docker compose up -d --build
TODAY'S RESULTS  less data/results/$(date +%F).jsonl
EMERGENCY        docker compose down   +   revoke the API key in the CRM

FILES
  .env           API key and switches         (section 7)
  config.json    queries, seeds, limits       (section 6)
  data/          local results and outbox     (section 10)

IN THE CRM
  Settings → API Keys      create / revoke the key (preset "Scraper")
  Settings → Data Sources  scraper activity (source "lead-scraper")
  Leads, filter Source     leads the scraper created
  Companies                HR contacts per company
↑ Top