yacy_search_server/README.md

200 lines
7.6 KiB
Markdown
Raw Normal View History

# YaCy
[![Gitter](https://badges.gitter.im/yacy/yacy_search_server.svg)](https://gitter.im/yacy/yacy_search_server?utm_source=badge&utm_medium=badge&utm_campaign=pr-badge&utm_content=badge)
[![Build Status](https://travis-ci.org/yacy/yacy_search_server.svg?branch=master)](https://travis-ci.com/yacy/yacy_search_server)
[![Deploy](https://www.herokucdn.com/deploy/button.svg)](https://heroku.com/deploy)
## What is this?
2020-02-06 14:49:31 +01:00
The YaCy search engine software provides results from a network of independent peers,
instead of a central server. It is a distributed network where no single entity decides
what to list or order it appears in.
User privacy is central to YaCy, and it runs on each user's computer, where search terms are
hashed before they being sent to the network. Everyone can create their individual
search indexes and rankings, and a truly customized search portal.
Each YaCy user is either part of a large search network (search indexes can be
exchanged with other installation over a built-in peer-to-peer network protocol)
or the user runs YaCy to produce a personal search portal that is either public or private.
YaCy search portals can also be placed in an intranet environment, making
it a replacement for commercial enterprise search solutions. A network
scanner makes it easy to discover all available HTTP, FTP and SMB servers.
To create a web index, YaCy has a web crawler for
2020-02-06 14:49:31 +01:00
everybody, free of censorship and central data retention:
- Search the web (automatically using all other YaCy peers)
- Co-operative crawling; support for other crawlers
- Intranet indexing and search
- Set up your own search portal
- All users have equal rights
- Comprehensive concept to anonymize the users' index
2020-02-06 14:49:31 +01:00
To be able to perform a search using the YaCy network, every user has to set up
their own node. More users means higher index capacity and better distributed
indexing performance.
## License
2021-03-11 12:23:53 +01:00
This project is available as open source under the terms of the GPL 2.0 Or later. However, some elements are being licensed under GNU Lesser General Public License. For accurate information, please check individual files. As well as for accurate information regarding copyrights.
2020-02-06 14:49:31 +01:00
The (GPLv2+) source code used to build YaCy is distributed with the package (in /source and /htroot).
## Where is the documentation?
- [Homepage](https://yacy.net)
2022-02-03 13:27:06 +01:00
- [International Forum](https://community.searchlab.eu)
- [German wiki](https://wiki.yacy.net/index.php/De:Start)
- [Esperanto wiki](https://wiki.yacy.net/index.php/Eo:Start)
- [French wiki](https://wiki.yacy.net/index.php/Fr:Start)
- [Spanish wiki](https://wiki.yacy.net/index.php/Es:Start)
- [Russian wiki](https://wiki.yacy.net/index.php/Ru:Start)
- [Video tutorials in English](https://yacy.net/en/Tutorials.html) and [video tutorials in German](https://yacy.net/de/Lehrfilme.html)
2020-02-06 14:49:31 +01:00
All these have (YaCy) search functionality combining all these locations into one search result.
## Dependencies? What other software do I need?
You need Java 1.8 or later to run YaCy. (No Apache, Tomcat or MySQL or anything else)
2020-02-06 14:49:31 +01:00
YaCy also runs on IcedTea 3.
See https://icedtea.classpath.org
2020-02-06 14:49:31 +01:00
## Start and stop it
2020-02-06 14:49:31 +01:00
Startup and shutdown:
2020-02-06 14:49:31 +01:00
- GNU/Linux and OpenBSD:
- Start by running `./startYACY.sh`
- Stop by running `./stopYACY.sh`
2020-02-06 14:49:31 +01:00
- Windows:
- Start by double-clicking `startYACY.bat`
- Stop by double-clicking `stopYACY.bat`
2020-02-06 14:49:31 +01:00
- macOS:
Please use the Mac app and start or stop it like any
other program (double-click to start)
2020-02-06 14:49:31 +01:00
## The administration interface
2020-02-06 14:49:31 +01:00
A web server us brought up after starting YaCy.
Open this URL in your web-browser:
http://localhost:8090
2020-02-06 14:49:31 +01:00
This presents you with the personal search and administration interface.
2020-02-06 14:49:31 +01:00
## (Headless) YaCy server installation
2020-02-06 14:49:31 +01:00
YaCy will authorize users automatically if they
access the server from its localhost. After about 10 minutes a random
password is generated, and then it is no longer possible to log in from
a remote location. If you install YaCy on a server that is not your
2020-02-06 14:49:31 +01:00
workstation you must set an admin account immediately after the first start-up.
Open:
http://<remote-server-address>:8090/ConfigAccounts_p.html
2020-02-06 14:49:31 +01:00
and set an admin account.
2020-02-06 14:49:31 +01:00
## YaCy in a virtual machine or a container
2020-02-06 14:49:31 +01:00
Use virtualization software like VirtualBox or VMware.
The following container technologies can deploy locally, on remote machines you own, or in the 'cloud' using a provider by clicking "Deploy" at the top of the page:
2016-07-13 01:06:33 +02:00
### Docker
2020-02-06 14:49:31 +01:00
More details in the [docker/Readme.md](docker/Readme.md).
2016-07-13 01:06:33 +02:00
2020-02-06 14:49:31 +01:00
### [Heroku](https://www.heroku.com/)
2020-02-06 14:49:31 +01:00
PaaS (Platform as a service)
More details in [Heroku.md](Heroku.md).
## Port 8090 is bad, people are not allowed to access that port
You can forward port 80 to 8090 with iptables:
2018-11-05 08:27:17 +01:00
```bash
iptables -t nat -A PREROUTING -p tcp --dport 80 -j REDIRECT --to-port 8090
2018-11-05 08:27:17 +01:00
```
On some operating systems, access to the ports you are using must be granted first:
2018-11-05 08:27:17 +01:00
```bash
iptables -I INPUT -m tcp -p tcp --dport 8090 -j ACCEPT
2018-11-05 08:27:17 +01:00
```
2020-02-06 14:49:31 +01:00
## Scaling, RAM and disk space
2020-02-06 14:49:31 +01:00
You can have many millions web pages in your own search index.
By default, 600MB RAM is available to the Java process.
2020-02-06 14:49:31 +01:00
The GC process will free the memory once in a while. If you have less than
100000 pages you could try 200MB till you hit 1 million.
[Here](http://localhost:8090/Performance_p.html) you can adjust it.
Several million web pages may use several GB of disk space, but you can
adjust it [Here](http://localhost:8090/ConfigHTCache_p.html) to fit your needs.
2020-02-06 14:49:31 +01:00
## Help develop YaCy
2020-02-06 14:49:31 +01:00
Join the large number of contributors that make YaCy what it is;
community software.
To start developing YaCy in **Eclipse**:
- Clone https://github.com/yacy/yacy_search_server.git using build-in Eclipse features (File -> Import -> Git)
- or Download source form this side (download button "Code" -> download as Zip -> and unpack)
- Import a Gradle project (File -> Import -> Gradle -> Existing Gradle Project).
- in the tab "Gradle Tasks" are tasks available to use build the project (e.g. build -> build or application -> run)
To start developing YaCy in **Netbeans**:
- clone https://github.com/yacy/yacy_search_server.git (Team → Git → Clone)
- if you checked "scan for project" you'll be asked to open the project
- Open the project (File → Open Project)
- you may directly use all the Netbeans build feature.
2022-02-03 13:27:06 +01:00
To join our development community, got to https://community.searchlab.eu
2020-02-06 14:49:31 +01:00
Send pull requests to https://github.com/yacy/yacy_search_server
2020-02-06 14:49:31 +01:00
## Compile from source
2020-02-06 14:49:31 +01:00
The source code is bundled with every YaCy release. You can also get YaCy
from https://github.com/yacy/yacy_search_server by cloning the repository.
```
git clone https://github.com/yacy/yacy_search_server
```
Compiling YaCy:
- You need Java 1.8 and ant
- See `ant -p` for the available ant targets
2020-02-06 14:49:31 +01:00
## APIs and attaching software
2020-02-06 14:49:31 +01:00
YaCy has many built-in interfaces, and they are all based on HTTP/XML and
HTTP/JSON. You can discover these interfaces if you notice the orange "API" icon in
the upper right corner of some web pages in the YaCy web interface. Click it, and
2020-02-06 14:49:31 +01:00
you will see the XML/JSON version of the respective webpage.
You can also use the shell script provided in the /bin subdirectory.
The shell scripts also calls the YaCy web interface. By cloning some of those
scripts you can easily create more shell API access methods.
## Contact
2022-02-03 13:27:06 +01:00
[Visit the international YaCy forum](https://community.searchlab.eu)
2020-02-06 14:49:31 +01:00
where you can start a discussion there in your own language.
2020-02-06 14:49:31 +01:00
Questions and requests for paid customization and integration into enterprise solutions.
can be sent to the maintainer, Michael Christen per e-mail (at mc@yacy.net)
with a meaningful subject including the word 'YaCy' to prevent it getting stuck in the spam filter.
2020-02-06 14:49:31 +01:00
- Michael Peter Christen