headless-chrome-crawler

(β˜… 5,636)

Distributed crawler powered by Headless Chrome

  • .editorconfig
  • .eslintrc.js
  • .gitignore
  • commitlint.config.js
  • Dockerfile
  • index.js
  • LICENSE
  • package.json
  • README.md
  • tsconfig.json
  • yarn.lock

# Installation Guide

1. Get the code
git clone https://github.com/yujiosaka/headless-chrome-crawler

Downloads the entire project code from GitHub to your computer.

cd headless-chrome-crawler

Moves into the project folder you just downloaded.

2. Docker

Easy Recommended
Prerequisites
  • Git Needed to download the project code from GitHub.
  • Docker Desktop Needed to build and run containers. Install it and keep it running in the background.
docker build -t headless-chrome-crawler .

Builds a runnable image based on the Dockerfile.

docker run -p 8080:80 headless-chrome-crawler

Runs the built image as an actual container.

βœ… Run docker compose ps to check the containers are Up. If the README mentions a port, open http://localhost:PORT in your browser.

3. Node.js

Easy
Prerequisites
  • Git Needed to download the project code from GitHub.
  • Node.js Node.js must be installed to use npm. The LTS version is recommended.
yarn add headless-chrome-crawler

Type this command into your terminal and run it.

βœ… After running the command, open the address shown in the terminal (usually something like http://localhost:3000) in your browser.

Pulled directly from this repo's README.

// repository documentation