#Eroforze Lead Scraper: Operator Manual
This manual covers everything needed to install, run, configure, monitor and troubleshoot the lead scraper. You don't need to read the code.
#1. What the scraper does
The scraper is a background service. It runs in cycles, by default one every 60 minutes. Each cycle has two phases.
#Phase 1: Discovery (new companies)
- It searches the web for each query in
config.json, for example "software development company in Chennai". - It also adds any seed domains you list yourself.
- It removes directories, job boards, social networks and marketplaces (LinkedIn, Justdial, Naukri, Clutch, Glassdoor and so on).
- It asks the CRM whether each remaining domain already exists. Existing companies are skipped.
- It visits each new website: the homepage plus up to 5 pages such as Contact, About, Team and Careers.
- If the site has a company name and at least one email or phone number, it uploads it to the CRM as a new lead, tagged with the query that found it.
#Phase 2: Enrichment and verification (existing leads)
- It reads every lead in the CRM through the API, oldest first.
- It picks leads that are missing something and haven't been checked in the last 30 days.
- For each one it:
- finds the website if the lead only has a company name,
- crawls the website for emails, phones, address, country, company LinkedIn and careers page,
- finds the person's LinkedIn if the lead is a person,
- finds HR or recruiting people at the company (team page first, then public LinkedIn search results),
- verifies the email address,
- checks that the website is still live.
- It sends only the missing fields back to the CRM. Nothing already in the CRM is ever overwritten, except the verification fields the scraper owns (email status, "verified", website status).
- Each HR person found becomes a separate lead linked to the company, and appears as an HR contact on the Companies page.
#Where data goes
Every result is written both to the CRM and to files on the server (data/). If the CRM is unreachable, uploads wait in a local outbox and are resent automatically on the next cycle, so nothing is lost.
#What it does not do
- It does not use AI, paid data providers or paid APIs.
- It never logs in to LinkedIn or scrapes LinkedIn directly. It only reads public search-engine results.
- It never sends emails. Even the optional mailbox check stops before a message is sent.
#2. Before you start
#Server requirements
| Item | Requirement |
|---|---|
| Operating system | Any 64-bit Linux (Ubuntu 22.04/24.04 recommended). macOS works for testing. |
| CPU / RAM | 1 vCPU and 1 GB RAM minimum; 2 vCPU / 2 GB recommended. |
| Disk | 2 GB free (the Docker images plus result files). |
| Software | Docker with Docker Compose v2 (recommended), or Node.js 22.18 or newer. |
| Network out | Port 443 (CRM API and websites) and 80 (some websites). Port 25 only if you enable mailbox checks. |
| Network in | None. The scraper opens no ports. |
Install Docker on Ubuntu if it isn't there:
curl -fsSL https://get.docker.com | sudo sh
sudo usermod -aG docker $USER # then log out and back in
docker compose version # should print v2.x
#What you need from the CRM
An API key with read and write access to leads. Section 3.1 explains how to create one.
#3. Installation
#3.1 Create the API key in the CRM
- Sign in to the CRM and open Settings → API Keys.
- In Generate API Key, fill in:
- Key name: something you'll recognise, for example
Lead scraper (server 1). - Source label:
lead-scraper. Use the same value asSCRAPER_SOURCEin.env(the default islead-scraper). - Product: the product new leads should belong to. Leave it as Default if unsure.
- Access: choose Scraper ("Send leads and check for duplicates").
- Key name: something you'll recognise, for example
- Click the generate button. The CRM shows "Copy your new key now". Copy the key (it starts with
efz_). It's shown only once. If you lose it, revoke it and create a new one.
One key per server is best practice. You can see each key's activity and revoke it on its own.
#3.2 Copy the folder to the server
From the machine that has the scraper folder:
rsync -av --exclude node_modules --exclude data scraper/ user@your-server:~/scraper/
# or: scp -r scraper user@your-server:~/
#3.3 Run the installer
ssh user@your-server
cd ~/scraper
./setup.sh
The installer:
- creates
.envfrom.env.exampleif it doesn't exist, - asks you to paste the API key and saves it in
.env, - generates a random secret for the private search engine (SearXNG),
- builds the scraper image and starts two containers that restart automatically, including after a server reboot:
eroforze-scraper: the scraper itself,eroforze-searxng: a private search engine only the scraper can reach,
- runs a connection check and prints the result.
A healthy check looks like this:
OK CRM API: https://site-crm-data.eroforze.com accepts key "Lead scraper (server 1)"
OK SearXNG: http://searxng:8080 answers JSON searches
OK Web search: 10 results
OK DNS (MX lookups): 5 mail servers for gmail.com
OK Outbound port 25: blocked (fine, SMTP_VERIFY is off)
If anything shows FAIL, see Troubleshooting. The scraper keeps running either way.
#3.4 Installing without Docker
If Docker isn't available, setup.sh installs with Node.js instead:
- Run as root: it installs a systemd service called
eroforze-scraperthat starts on boot. - Run as a normal user: it starts the scraper in the background and writes logs to
data/scraper.log.
Without Docker there's no private search engine, so search falls back to public DuckDuckGo. That allows only a handful of searches per hour, which limits discovery, HR lookups and LinkedIn lookups. Use Docker for production.
#4. Your first run (safe test)
Before letting the scraper write to the CRM, do one dry run. It crawls and saves everything locally but sends nothing to the CRM.
cd ~/scraper
docker compose run --rm -e RUN_ONCE=true -e DRY_RUN=true scraper
When it finishes, inspect what it would have uploaded:
ls data/results/
less data/results/$(date +%F).jsonl
Each line is one record:
"type":"new-company": a company discovery would add,"type":"enrich": the fields it would fill on an existing lead (filled,patch),"type":"hr-contact": an HR person lead it would create,"type":"cycle": the summary of the run.
When you're happy, start it for real:
docker compose up -d
docker compose logs -f scraper
Tip: for the first live run, set
"maxNewCompaniesPerCycle": 10and"maxLeadsPerCycle": 25inconfig.jsonso you can review a small batch in the CRM first.
#5. Day-to-day commands
Run all commands from the scraper folder.
#With Docker (normal setup)
| Task | Command |
|---|---|
| Watch live logs | docker compose logs -f scraper |
| Last 200 log lines | docker compose logs --tail 200 scraper |
| Is it running? | docker compose ps |
| Stop everything | docker compose down |
| Start again | docker compose up -d |
Apply changes to config.json or .env | docker compose up -d --force-recreate scraper |
| Run one cycle now and exit | docker compose run --rm -e RUN_ONCE=true scraper |
| Dry run (nothing sent to the CRM) | docker compose run --rm -e RUN_ONCE=true -e DRY_RUN=true scraper |
| Connection check | docker compose run --rm scraper node src/check.ts |
| Search-engine logs | docker compose logs --tail 100 searxng |
Stopping is graceful. On docker compose down or restart, the scraper finishes the lead it's working on and then exits. Anything not yet sent is either already in the CRM or in the outbox.
#Without Docker (systemd)
| Task | Command |
|---|---|
| Live logs | journalctl -u eroforze-scraper -f |
| Status | systemctl status eroforze-scraper |
| Stop / start / restart | sudo systemctl stop|start|restart eroforze-scraper |
| One cycle now | RUN_ONCE=true node src/index.ts |
| Dry run | npm run dry-run |
| Connection check | npm run check |
#6. Choosing what it finds (config.json)
config.json controls what the scraper looks for and how much it does per cycle. After editing, apply the change with docker compose up -d --force-recreate scraper.
The file must stay valid JSON: double quotes, no trailing commas, no comments. If unsure, paste it into any online "JSON validator" first.
#6.1 Discovery queries
"discovery": {
"enabled": true,
"maxNewCompaniesPerCycle": 50,
"revisitDays": 90,
"queries": [
{
"query": "software development company",
"locations": ["Bangalore", "Chennai", "Hyderabad"],
"country": "IN",
"industry": "Software Development",
"tags": ["it-services"],
"product": "edgeryt-hire",
"pages": 2
}
],
"seeds": [],
"blocklist": []
}
| Field | What it does |
|---|---|
query | The search phrase. Describe the type of company, for example "logistics company", "manufacturing company hiring" or "fintech startup". |
locations | Each location is searched separately as "query in location". Leave it out or use [] to search the query alone. |
country | ISO code (IN, AE, US, GB, SG…) used only if the website doesn't reveal its country. |
industry | Saved in the lead's Industry field. |
tags | Added to customFields.tags together with discovered. Use them to filter or build campaigns later. |
product | CRM product key for these leads. If omitted, SCRAPER_PRODUCT from .env is used, then the API key's product. |
pages | How many pages of search results to read per location (about 10 results per page). Start with 1–2. |
Other discovery settings:
| Setting | Meaning |
|---|---|
maxNewCompaniesPerCycle | Upper limit of new companies added per cycle. |
revisitDays | A domain discovery has already looked at is ignored for this many days. |
enabled | false turns discovery off; only enrichment runs. |
How queries are worked through. Each cycle goes through the queries in order until it has enough candidates. Search results are cached for 14 days and checked domains are remembered for 90 days, so later cycles naturally move on to queries and locations not yet covered. Long query lists are fine; they're simply spread over several cycles.
Writing good queries:
- Be specific about the company type:
"SaaS company"beats"companies". - Add intent words that suit your product:
"hiring","startup","careers". - Use cities rather than whole countries for better local results.
- Avoid queries that mainly return lists or articles, such as
"top 10 IT companies". These pages are filtered out, so they waste searches.
#6.2 Seed domains
Seeds are companies you already know you want added and enriched, such as event attendees or a purchased list of websites:
"seeds": [
{ "domain": "acme.com", "country": "IN", "industry": "Manufacturing", "tags": ["expo-2026"] },
{ "domain": "https://www.example.ae", "country": "AE", "tags": ["partner-referral"] }
]
Seeds are processed before search queries. They go through the same checks: skipped if already in the CRM, and added only if the site shows a name and a way to contact the company.
#6.3 Blocklist
The scraper already ignores well-known directories, job boards, social networks, news sites and marketplaces. Add anything else that keeps appearing but isn't a real prospect:
"blocklist": ["competitor.com", "somedirectory.in", "examplelistings"]
An entry with a dot blocks that exact domain. An entry without a dot blocks every domain containing that name ("examplelistings" blocks examplelistings.com, examplelistings.co.in and so on).
#6.4 Enrichment settings
"enrichment": {
"enabled": true,
"maxLeadsPerCycle": 200,
"refreshDays": 30,
"product": "",
"source": "",
"fillCompanyEmail": true,
"findHrContacts": true,
"createHrLeads": true,
"maxHrContactsPerCompany": 2,
"findPersonLinkedin": true,
"findMissingWebsites": true
}
| Setting | Meaning |
|---|---|
maxLeadsPerCycle | How many leads are enriched per cycle. |
refreshDays | After a lead is checked, it's left alone for this many days, even if nothing was found. Emails are also re-verified after this period. |
product | Only enrich leads of this product (key or name). Empty means all products. |
source | Only enrich leads from this source, for example csv-import. Empty means all sources. |
fillCompanyEmail | For company leads (no person name) with no email, use the best company inbox as the email: HR inbox first, then a general inbox such as info@. |
findHrContacts | Look for HR and recruiting people at each company. |
createHrLeads | Create a separate lead for each HR person found. If false, only the company lead's hr_name / hr_title fields are filled. |
maxHrContactsPerCompany | Maximum HR people added per company. |
findPersonLinkedin | For person leads without a LinkedIn URL, search for their profile. |
findMissingWebsites | For leads with a company name but no website, search for the official website. |
Leads with status Archived or Lost are never enriched.
#6.5 Crawl settings
"crawl": {
"maxPagesPerSite": 6,
"maxConcurrency": 10,
"maxRequestsPerMinute": 120,
"respectRobotsTxt": true,
"timeoutSecs": 30
}
| Setting | Meaning |
|---|---|
maxPagesPerSite | Pages read per website, homepage included. Only contact, about, team, careers, leadership and imprint-style pages are followed. |
maxConcurrency | Websites fetched in parallel. |
maxRequestsPerMinute | Overall page-fetch speed limit. Each site is also limited to one page per second. |
respectRobotsTxt | Obeys each site's robots.txt. Keep this true. |
timeoutSecs | A page that takes longer than this is skipped. |
#7. Settings reference (.env)
.env holds the API key and switches. After editing, apply with docker compose up -d --force-recreate scraper.
| Setting | Default | Meaning |
|---|---|---|
EROFORZE_API_KEY | (none) | Required. The CRM API key (efz_…). |
EROFORZE_API_URL | https://site-crm-data.eroforze.com | CRM API address. |
SCRAPER_SOURCE | lead-scraper | Source label on every lead the scraper creates. Shown in Settings → Data Sources and usable as a filter. |
SCRAPER_PRODUCT | (empty) | Product key for new leads. Empty means the API key's product, then the workspace default. |
CYCLE_INTERVAL_MINUTES | 60 | Pause between the end of one cycle and the start of the next. |
RUN_ONCE | false | true runs one cycle and exits (for cron or testing). |
DRY_RUN | false | true saves locally but sends nothing to the CRM. |
SEARCH_ENABLED | true | false disables all web searches (no discovery by search, no HR or LinkedIn lookups, no missing-website lookups). Seeds and website crawling still work. |
SEARCH_ENGINES | searxng,duckduckgo | Tried in order; falls back to the next one if the first fails or rate limits. |
SEARXNG_URL | set by Docker | Address of the private search engine. Leave empty in .env when using Docker. |
SEARCH_DELAY_MS | 8000 | Minimum gap between searches on public DuckDuckGo. SearXNG uses a third of this. Raise it if searches get blocked. |
SEARCH_CACHE_DAYS | 14 | How long search results are reused. |
SEARXNG_SECRET | generated | Secret for the private search engine. Don't share it or change it while running. |
SMTP_VERIFY | false | true turns on mailbox-level email checks and HR email guessing. See section 11. |
SMTP_HELO_DOMAIN | localhost | Domain the scraper introduces itself with to mail servers. Use a real domain you own. |
SMTP_FROM | verify@localhost | Sender address used in the check (no email is ever sent). Use an address on that domain. |
SMTP_TIMEOUT_MS | 12000 | How long to wait for a mail server. |
PROXY_URLS | (empty) | Comma-separated proxies for website crawling, for example http://user:pass@1.2.3.4:8000. |
USER_AGENT | a desktop Chrome | Browser identity sent to websites. |
LOG_LEVEL | info | debug shows every search, crawl failure and SMTP reply. |
DATA_DIR | data | Where local files are kept. |
#8. Seeing the results in the CRM
#8.1 Data Sources
Settings → Data Sources shows a row for the scraper's source label (lead-scraper). It lists the number of requests, leads received, New, Updated and the last upload time. Use it to confirm the scraper is alive.
#8.2 Finding scraper leads
- Leads page, filter Source =
lead-scraper: every company and HR person the scraper created. - Leads that already existed keep their original source. The scraper only fills their gaps, and the changes appear in each lead's audit log under the API key's name.
#8.3 Fields the scraper fills
Standard fields (filled only if empty): Email, Phone, Website, Company, Country, City, State, LinkedIn URL. Email status and Verified at are set by verification.
Custom fields on company leads and existing leads:
| Custom field | Example | Meaning |
|---|---|---|
hr_email | careers@acme.com | HR or recruitment inbox found on the website. |
hr_name, hr_title | Priya Sharma, Senior HR Business Partner | HR person found for the company. |
company_email | info@acme.com | General inbox. |
company_linkedin | https://www.linkedin.com/company/acme | Company LinkedIn page. |
careers_url | https://acme.com/careers | Careers or jobs page. |
tags | discovered, it-services | Tags from the discovery query. |
lead_type | company / person | What kind of lead the scraper created. |
discovered_via | search: software development company in Chennai | The query or seed that found it. |
discovered_at | ISO date | When it was discovered. |
search_location | Chennai | The search location that found it. |
country_source | website address, domain, phone code, page language, search location | Where the country came from. |
country_detected | AE | The website's address says a different country from the one in the CRM. Worth a manual look. |
verified | true / false | Email confirmed good or bad. Shown as Email verified on the Companies page. |
email_check | listed on the company website, mail server found | Why the email got its status. |
email_found_by | name pattern + mail server check | The email was guessed from the person's name and confirmed by their mail server. |
website_status | live / unreachable | Whether the website loaded. |
enriched_at | ISO date | When the scraper last checked this lead. |
enrich_result | found phone, country; still missing hr_email | What it found and what's still missing. |
HR person leads created by the scraper have Full name, Title, LinkedIn URL, Company, Website and Country, plus:
| Custom field | Meaning |
|---|---|
linkedin_role | hr or recruiting. |
linkedin_profile_of | "Name - Title" (used by the Companies page). |
hr_linkedin / recruiting_linkedin | Their LinkedIn profile. |
parent_lead | Lead ID of the company lead they were found for. |
tags | hr-contact, scraped. |
found_via | company website or linkedin search. |
#8.4 Companies page
Open a company on the Companies page to see its contacts. HR people the scraper created appear at the top as named HR contacts with their LinkedIn, together with the HR inbox (hr_email) and the General inbox.
#8.5 Building campaigns from scraper leads
Use the audience filters (Source = lead-scraper, Industry, Country, Product). Only leads with a usable email and the right consent status are emailed, so check consent rules before sending to cold leads.
#9. Reading the logs
Example of a normal cycle:
INFO Eroforze lead scraper starting {"api":"https://site-crm-data.eroforze.com","source":"lead-scraper","search":["searxng","duckduckgo"],"smtpVerify":false,"dataDir":"/app/data"}
INFO Connected to the CRM as API key "Lead scraper (server 1)" (default product edgeryt-hire)
INFO Search "software development company in Chennai": 8 website(s)
INFO Crawling 8 new website(s)
INFO Uploaded 5 lead(s): 5 new, 0 updated, 0 unchanged, 0 rejected
INFO New company Creatah Software Technologies (creatah.com) career@creatah.com IN
INFO Discovery done {"searched":1,"candidates":8,"alreadyInCrm":0,"notACompany":3,"created":5}
INFO Crawling 20 website(s)
INFO LD-7K3QX9MZ Freshworks: found phone, city, state, country, careers_url, company_linkedin, email, hr_name, hr_title; still missing hr_email
INFO LD-2HF8PL0Q Dead Co: nothing new found; still missing phone, country, linkedin, hr_email, hr_contact
INFO Uploaded 3 lead(s): 3 new, 0 updated, 0 unchanged, 0 rejected
INFO Enrichment done {"checked":20,"updated":12,"fieldsFilled":64,"verified":15,"hrLeads":3}
INFO Cycle finished {"discovery":{...},"enrichment":{...},"searches":42,"websitesCrawled":28,"minutes":6.4}
INFO Next cycle in 60 minute(s)
What the summary numbers mean:
| Number | Meaning |
|---|---|
searched | Discovery searches run. |
candidates | New domains worth crawling (not already in the CRM). |
alreadyInCrm | Domains skipped because the company is already a lead. |
notACompany | Websites skipped: unreachable, or no name or contact details. |
created | New company leads added. |
checked | Existing leads looked at. |
updated | Leads that gained at least one field. |
fieldsFilled | Total fields filled. |
verified | Emails that got a verification status. |
hrLeads | HR person leads created. |
searches | Web searches actually sent this cycle (cached ones aren't counted). |
Warnings you may see, and what they mean:
| Message | Meaning | Action |
|---|---|---|
duckduckgo is rate limiting us; pausing it for 30 minutes | The public search engine is blocking us for now. | Normal now and then. If constant, make sure SearXNG is running (Docker setup). |
searxng is rate limiting us… | Every engine inside SearXNG is blocked. | Raise SEARCH_DELAY_MS, reduce queries or pages. |
… failed after 6 attempts … Saved … to the outbox | The CRM couldn't be reached after retries. | Nothing to do; it's resent next cycle. Check the CRM if it repeats. |
Update of LD-… refused (400): … | The CRM rejected a change (for example an invalid value). | Read the message; usually a single bad value. |
Lead rejected: … | One new lead failed validation (for example a malformed email). | Usually harmless; the rest of the batch was saved. |
API key rejected (401) / revoked | The key is wrong or was revoked. | Create a new key and update .env (section 12.3). |
missing leads:read, leads:write | The key lacks permissions. | Create a key with the Scraper or Full access preset. |
#10. Files saved on the server
Everything lives in scraper/data/:
| File | Contents | Safe to delete? |
|---|---|---|
results/YYYY-MM-DD.jsonl | Every new company, HR contact and enrichment (what was found, what was sent, whether the CRM accepted it), plus one summary line per cycle. | Yes. It's your local history and backup. Keep or archive as you like. |
outbox.jsonl | Uploads waiting for the CRM. Normally empty. | No. Deleting it loses unsent data. |
seen-domains.json | Domains discovery has already checked, with dates. | Yes, if you want discovery to re-check every domain. |
search-cache.json | Recent search results. | Yes. Searches will simply run again. |
scraper.log, scraper.pid | Only in no-Docker mode. | Yes, when stopped. |
Useful one-liners:
# today's new companies
grep '"type":"new-company"' data/results/$(date +%F).jsonl | wc -l
# everything sent about one lead
grep 'LD-7K3QX9MZ' data/results/*.jsonl
# cycle summaries for today
grep '"type":"cycle"' data/results/$(date +%F).jsonl
# export today's new companies as CSV (needs jq)
grep '"type":"new-company"' data/results/$(date +%F).jsonl \
| jq -r '.lead | [.companyName, .website, .email, .phone, .country, .industry] | @csv'
#11. Email verification explained
Every email the scraper adds, and every existing email not verified in the last refreshDays, gets a status in the CRM's Email status field.
| Status | Meaning | Safe to email? |
|---|---|---|
valid | The domain receives mail and either the company publishes the address on its own website, or the mail server confirmed the mailbox (SMTP mode). | Yes. |
catch_all | The company's mail server accepts any address, so this one can't be confirmed (SMTP mode only). | Probably; bounces are possible. |
risky | Disposable email domain, or the mail server temporarily refused the check. | Avoid or check by hand. |
invalid | Bad format, the domain doesn't exist or has no mail server, or the mail server said the mailbox doesn't exist. | No. |
unknown | The domain receives mail, but the specific mailbox couldn't be confirmed. | Use with care. |
Two rules protect existing data:
- An
unknownresult never replaces a status someone already set. customFields.verifiedis set totrueonly forvalidand tofalseonly forinvalid.
#Turning on mailbox checks (SMTP_VERIFY=true)
Turning this on adds three things:
- a check with the company's mail server that the mailbox really exists,
- detection of catch-all domains,
- guessing personal and HR emails from a person's name (for example
priya.sharma@,priya@,psharma@), accepted only when the mail server confirms the address and the domain isn't catch-all.
Requirements:
- Outbound port 25 must be open. Run the connection check; it reports "Outbound port 25: open/blocked". AWS, Google Cloud, Azure and many VPS providers block it by default. Some unblock it on request.
- Set
SMTP_HELO_DOMAINandSMTP_FROMto a real domain you control, ideally one whose server IP has a matching reverse-DNS (PTR) record. Without that, some mail servers refuse to answer. - Don't use your main email-sending server's IP. Heavy checking can affect that IP's reputation.
#12. Maintenance
#12.1 Changing settings
- Edit
config.jsonor.env. - Run
docker compose up -d --force-recreate scraper. - Watch
docker compose logs -f scraperfor the "starting" line with your new settings.
#12.2 Updating the scraper code
Copy the new src/ (and package.json / package-lock.json if they changed), then rebuild:
cd ~/scraper
docker compose up -d --build
Your .env, config.json and data/ are kept.
To update the search engine image now and then:
docker compose pull searxng && docker compose up -d searxng
#12.3 Rotating or replacing the API key
- In the CRM, Settings → API Keys, create a new key (preset Scraper).
- Put it in
.envasEROFORZE_API_KEY=…. - Run
docker compose up -d --force-recreate scraperand confirm "Connected to the CRM as API key …" in the logs. - Revoke the old key in the CRM.
#12.4 Making the scraper re-check specific leads
A lead is skipped for refreshDays after its last check. To re-check it sooner, clear its enriched_at custom field, either in the CRM or through the API:
curl -X PATCH "https://site-crm-data.eroforze.com/api/v1/leads/LD-7K3QX9MZ" \
-H "Authorization: Bearer $EROFORZE_API_KEY" -H "Content-Type: application/json" \
-d '{"customFields":{"enriched_at":null}}'
To re-check everything sooner, lower refreshDays temporarily (for example to 1), then set it back.
#12.5 Backups and moving to another server
- Back up
scraper/.env,scraper/config.jsonand, optionally,scraper/data/. - To move: copy the whole folder (without
node_modules), run./setup.shon the new server, and stop the old one withdocker compose down. Never run two copies with the same configuration at once. They'd do the same work twice. The CRM's duplicate protection prevents duplicate leads, but you'd waste searches.
#12.6 Running more than one scraper
To cover different markets in parallel, run separate copies of the folder on the same or different servers. Give each copy its own config.json queries, its own API key and, if you want them told apart in Data Sources, its own SCRAPER_SOURCE. Set enrichment.enabled to false on all but one copy, so only one enriches existing leads.
#12.7 Disk usage
Result files are small, typically a few MB per month. Container logs are capped at 50 MB for the scraper and 30 MB for the search engine. To remove old result files:
find data/results -name '*.jsonl' -mtime +180 -delete
#13. Troubleshooting
#The connection check fails
| Check | Likely cause | Fix |
|---|---|---|
CRM API FAIL 401 | Wrong or mistyped key. | Re-copy the key into .env, recreate the container. |
CRM API FAIL revoked | The key was revoked. | Create a new key (section 12.3). |
CRM API FAIL missing leads:read… | The key was created with the wrong access. | New key with the Scraper preset. |
CRM API FAIL network error | The server can't reach the CRM. | Test curl https://site-crm-data.eroforze.com/api/v1/health; check firewall and DNS. |
| SearXNG FAIL unreachable | The search container isn't running. | docker compose ps, then docker compose logs searxng, then docker compose up -d searxng. |
| SearXNG FAIL answered 403 | JSON output disabled. | Make sure searxng/settings.yml lists json under search.formats, then docker compose restart searxng. |
| Web search FAIL no results | All search engines are blocking the server's IP right now. | Wait 30–60 minutes; raise SEARCH_DELAY_MS; consider a server with a different IP. Crawling seeds and existing websites still works. |
| DNS FAIL | The server can't resolve domains. | Check /etc/resolv.conf or Docker DNS. |
| Port 25 FAIL | Port 25 blocked while SMTP_VERIFY=true. | Set SMTP_VERIFY=false, or ask your provider to unblock port 25. |
#The scraper runs but…
| Symptom | Cause and fix |
|---|---|
| No new companies | Check the logs for alreadyInCrm (the companies are already in the CRM) and notACompany (the sites have no contact details). Add more locations or queries, raise pages, or check searches aren't blocked. Domains are remembered for 90 days; delete data/seen-domains.json to re-check them. |
| "nothing new found" on most leads | The leads' websites don't publish the missing data, or the sites are built entirely in JavaScript (the crawler reads plain HTML). Normal for a portion of leads. |
| Few HR contacts | HR lookups depend on search. Make sure SearXNG works (connection check). Small companies often have no public HR profiles. |
| Few personal emails | Without SMTP_VERIFY=true, personal emails are only added when published on the company's site. See section 11. |
| Wrong country on some leads | When a website has no address, the country comes from the domain (.in → India), the phone code, or the query's country. Look at country_source on the lead. Existing countries are never changed; disagreements are flagged in country_detected. |
| Wrong company name | It's taken from the website's structured data, its site name or the page title. Fix it in the CRM; the scraper never overwrites it. |
| Leads in the wrong product | Set product on the query, or SCRAPER_PRODUCT, or the API key's product. |
| Container keeps restarting | docker compose logs --tail 50 scraper. Usually a missing or invalid API key, or invalid JSON in config.json. |
outbox.jsonl keeps growing | The CRM has been unreachable or failing for a while. Check the CRM and the API URL. Items are resent automatically once it's back. |
| High CPU or memory | Lower crawl.maxConcurrency (for example to 5) and maxRequestsPerMinute. |
#Getting more detail
Set LOG_LEVEL=debug in .env and recreate the container. You'll see every search query and result count, every page that failed to load, the reasons engines were skipped, and SMTP replies. Set it back to info afterwards; debug output is large.
#14. Safety, privacy and stopping in an emergency
#Stop immediately
docker compose down
To cut off CRM access completely, even if someone restarts the containers, revoke the API key in Settings → API Keys.
#Undoing a bad batch
- Every lead the scraper created has the scraper's source label, so you can find and bulk-review them on the Leads page (filter by Source).
- Every change to an existing lead is recorded in that lead's audit log, with the old and new values, under the API key's name.
data/results/*.jsonllists exactly what was sent and when.
#Good practice
- Consent and law: scraped contacts are cold leads. Follow the rules that apply to you (India's DPDP Act, GDPR for EU contacts, CAN-SPAM for the US, UAE PDPL). Always include an unsubscribe option; the CRM's campaigns already do.
- Be polite to websites: keep
respectRobotsTxt: true. Don't raise speed limits much beyond the defaults. - Search engines: automated searching can break search engines' terms. The private SearXNG plus caching keeps volume low, but large query lists mean more searches.
- LinkedIn: the scraper never logs in to or crawls LinkedIn. Profile links come only from public search results and companies' own websites, and are accepted only when the result names both the person and the company.
- Secrets:
.envcontains the API key. Don't commit it to git or share it. The folder's.gitignorealready excludes it.
#15. Speed and capacity
These are rough figures for the default settings on a 2 vCPU server. Actual numbers depend on how fast websites respond and how often search engines limit you.
| Activity | Typical speed |
|---|---|
| Website crawling | About 20 websites in 20–60 seconds (6 pages each, 10 in parallel). |
| Search (SearXNG) | One search every 3–4 seconds, about 1,000 per hour at most. Cached results are free. |
| Search (DuckDuckGo only, no Docker) | A few searches before a 30-minute pause. |
| Enrichment | 1–3 searches per lead (website, LinkedIn, HR) plus a crawl. Often 100–300 leads per hour. |
| Discovery | Each query and location costs pages searches, then a crawl of the new domains. |
Tuning tips:
- More leads per day: raise
maxLeadsPerCycleandmaxNewCompaniesPerCycle, or lowerCYCLE_INTERVAL_MINUTES. Watch the logs for rate-limit warnings. - Fewer search blocks: raise
SEARCH_DELAY_MS, lowerpages, or setfindPersonLinkedintofalse(it's the most search-hungry option). - Lighter on the server: lower
maxConcurrencyandmaxRequestsPerMinute.
#16. Quick reference card
INSTALL ./setup.sh
LOGS docker compose logs -f scraper
STATUS docker compose ps
CHECK docker compose run --rm scraper node src/check.ts
STOP / START docker compose down | docker compose up -d
APPLY CHANGES docker compose up -d --force-recreate scraper
RUN ONE CYCLE docker compose run --rm -e RUN_ONCE=true scraper
DRY RUN docker compose run --rm -e RUN_ONCE=true -e DRY_RUN=true scraper
UPDATE CODE docker compose up -d --build
TODAY'S RESULTS less data/results/$(date +%F).jsonl
EMERGENCY docker compose down + revoke the API key in the CRM
FILES
.env API key and switches (section 7)
config.json queries, seeds, limits (section 6)
data/ local results and outbox (section 10)
IN THE CRM
Settings → API Keys create / revoke the key (preset "Scraper")
Settings → Data Sources scraper activity (source "lead-scraper")
Leads, filter Source leads the scraper created
Companies HR contacts per company