yacy_search_server

mirror of https://github.com/yacy/yacy_search_server.git synced 2024-09-21 00:00:13 +02:00

Author	SHA1	Message	Date
luc	8c4ab9c76b	Added an option to eventually limit size of remote solr documents put to local index. See mantis #626.	2015-12-16 02:20:03 +01:00
luc	a2c08402af	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-12-15 23:30:30 +01:00
luc	70595d05d0	Modified MemoryControl.main() test to properly end for better results displaying.	2015-12-14 23:49:28 +01:00
sixcooler	1be67d9ab6	CachedSolrConnector was replaced by ConcurrentUpdateSolrConnector years ago - time to let it go Commented out unused table of cache-objects	2015-12-14 21:33:27 +01:00
reger	28b8bc290a	fix use of NETWORK_SEARCHVERIFY for rwi verification was not used to set the searchevent parameter (done in SearchEventCache.getEvent) - remove unused corresponding QueryParams.filterfailurls param.	2015-12-13 20:01:49 +01:00
reger	020630efd8	remove unused network scanner parameter from queryparameter Search event is not using networkscanner (removed filterscannerfail param always init to false)	2015-12-13 02:50:08 +01:00
luc	ad5586f8f6	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-12-08 03:35:36 +01:00
luc	8ebefa4233	Fixed MediaWiki import : DCEntry conversion to SolrInputDocument was failing. Looks like it was broken since Commit `b43811d38c`	2015-12-08 03:34:03 +01:00
luc	7736ee5a42	Updated MediaWimporter main() : display usage in console and stop properly without calling System.exit	2015-12-08 03:30:51 +01:00
reger	cdb8f3b10d	make current ranking score value avail. to search interface / api Update the result score result field with the result queue ranking value to reflect the actual calculated/used score, for rwi & solr stack results. (calc. etc. is unchanged, it's just that result entry carries the latest val as api retrieves the number from it)	2015-12-08 03:17:32 +01:00
luc	27d11f8671	Fixed isSolrDump function : PushBackInputStream was not unread when returning false (for example with a WikiMedia dump).	2015-12-07 21:58:36 +01:00
Michael Peter Christen	135a123a77	less logging in new language detection	2015-12-03 00:39:15 +01:00
Michael Peter Christen	ef8cd80593	fix for npe	2015-12-03 00:33:13 +01:00
reger	0347bfa71f	Apply collection query constraint/modifiert to rwi result stack. Collection is not available in pure rwi entries (but in local solr metadata) But if user wishes to filter by query constraint also rwi shall adhere to this (even if only rwi entries with parsed or solr received metadata may fit)	2015-12-02 22:57:59 +01:00
luc	2a67d2ba6f	Corrected error management for unsupported image formats, parsing errors, and unavailable resources : avoid logging to much Exceptions as these errors easily occur when searching images.	2015-12-01 01:06:01 +01:00
Michael Peter Christen	d6e9834040	Merge branch 'master' of https://github.com/Scarfmonster/yacy_search_server # Conflicts: # .classpath # build.xml	2015-11-30 16:54:54 +01:00
Michael Peter Christen	d82d311995	Merge branch 'master' of https://github.com/luccioman/yacy_search_server # Conflicts: # .classpath	2015-11-30 13:34:10 +01:00
reger	b5371ea8c1	read/init crawl queue in a thread to speed-up YaCy start on large existing crawler queues	2015-11-29 05:19:39 +01:00
reger	1160b13172	remove unused md5 from ViewFile servlet params	2015-11-28 23:09:15 +01:00
reger	e163ea88f6	fix vsdParser (Visio) parser return statement (final block un-necessary throw)	2015-11-28 02:43:38 +01:00
reger	b2c8bc0ae6	remove md5_s from default index fields it is not assigned a value / not used Due to above also excluded from transfer protocol.	2015-11-27 02:41:02 +01:00
luc	e40ae0943b	- No max dimensions specified : render raw image data when source and target image format are the same. - Corrected scaling condition.	2015-11-26 09:30:43 +01:00
reger	90686a75a2	fix flux factor (additional crawl delay by access count) calculation	2015-11-25 01:34:41 +01:00
luc	4af27289e5	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-23 09:01:25 +01:00
reger	297fdb60d3	throw exception if crawler hostqueue can't create hostpath directory. In rare cases hostname may not be a valid filesystem directory name, which can't be created (e.g. containing '*' char). To prevent crawl queue looping on this invalid entry by throwing a malformedurlexception.	2015-11-22 21:26:18 +01:00
luc	755efac17d	Use same max file size when loading all resource bytes or opening stream content	2015-11-20 19:35:39 +01:00
luc	bc6c79fc12	Corrected scaling function for non RGB images.	2015-11-20 14:35:36 +01:00
luc	1565559df8	Refactoring : extracted write InputStream method.	2015-11-20 09:42:24 +01:00
luc	f0478bb14d	BMP and ICO image formats support : integrated /haraldk/TwelveMonkeys imageio-bmp-3.2 library. - better BMP format flavours support - handle PNG encoded icons - handle transparency Added some javadoc url references to .classpath	2015-11-20 09:38:16 +01:00
luc	07437986e7	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-20 08:15:24 +01:00
reger	97cc03ef6a	start using a template for urlproxy header It is included as iframe /proxmsg/urlproxyheader.html to allow full servlet functionallity and flexibility to display some index/meta data in future.	2015-11-20 01:49:56 +01:00
luc	f01d49c37a	Process large or local file images dealing directly with content InputStream.	2015-11-18 10:15:38 +01:00
luc	3c4c77099d	If available, check content length before downloading. Check also content length is not over Integer.MAX_VALUE.	2015-11-18 10:11:38 +01:00
luc	5bbb2e1730	Ensure resource is closed when reading a full file InputStream	2015-11-18 10:08:06 +01:00
luc	6291a57300	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-18 08:49:31 +01:00
reger	0d3c5b223e	have psParser cleanup temp file	2015-11-17 23:45:29 +01:00
reger	7d0d19cb8e	avoid File.deleteOnExit() on temp files JVM registers each file in a list regardless of already deleted and never cleans up the list during runtime. This accumulates to a considerable amount of mem during large crawls and/or long uptime. To tackle this, all temp files are now created in a subdir of java.io.tmpdir and the jvm tmpdir property is set to this subdir, which is deleted by code on shutdown. Additionally let pdfParser use this tmp subdir too.	2015-11-17 22:27:07 +01:00
luc	bfe51001e3	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-17 08:30:32 +01:00
reger	02e4489a23	set tmpfile.deleteOnExit by default, to make sure files are removed on shutdown.	2015-11-16 21:37:45 +01:00
reger	2985baaa01	Exclude repetitive protocol part in tokenized url used as description if none is avail. from parser.	2015-11-16 01:06:20 +01:00
reger	ca3d26a401	harmonize wordsintitle & CollectionSchema.title_words_val calculation, remove obsolete partial init of wordreference from urimetadata	2015-11-15 06:06:37 +01:00
reger	52a9040ae6	Sort out double keywords (dc_subject) early in parsed documents - by direct using Set vs. List - remove not neede String[] getter	2015-11-13 01:48:28 +01:00
luc	49331dc523	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-12 08:21:56 +01:00
reger	47d70732f6	improve locale translator - skip empty line - robustness file section detection (space independant)	2015-11-11 00:57:51 +01:00
sixcooler	646afe9183	do not store subfield *_coordinate + make all num-fields being docvalues	2015-11-10 20:45:33 +01:00
sixcooler	194df613de	not using 'location' as defaultfacetfield - since we removed it being default.	2015-11-10 20:43:58 +01:00
sixcooler	d3b9349b6f	simplification / speedup of GenerationMemoryStrategy	2015-11-10 20:39:46 +01:00
sixcooler	4a905ec134	fix to not let the AccessTracker-Log grow to much, but have enough data to monitor. (+gitignore-correction)	2015-11-10 20:27:17 +01:00
reger	20e18d79f8	harmonize document title for archive parsers	2015-11-10 01:29:13 +01:00
luc	f11b5e8309	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-09 08:13:12 +01:00
reger	112ae013f4	update bzip and bzip parser process, to return one document for the file with combined parser results of the containing file and registers it with supplied url and mime of the archive.	2015-11-07 19:13:18 +01:00
reger	e76a90837b	update zip and tar parser process, to return one document for the file with combined parser results of the containing files.	2015-11-06 23:58:55 +01:00
luc	4e673ffc9a	Ensure closing of InputStream even when an exception occurs.	2015-11-05 09:40:24 +01:00
luc	10696b53f7	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-05 08:26:52 +01:00
reger	8532565c7d	optimize order of parsers to try - start with a parser matching the remote supplied mime	2015-11-04 21:52:02 +01:00
reger	681889ae64	use current tar library for untar files - remove old source copy	2015-11-04 02:57:00 +01:00
reger	5d71fc70e3	fix tarParser early exit on looping content - adjust check of data available according to doc - return null on no recognized content (to not exit TextParser next parser try) - use commons.compress directly	2015-11-03 22:14:14 +01:00
luc	bcc2e7cb5b	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-03 09:29:57 +01:00
reger	2fcf6f104c	fix bzipParser recognition - Bzip2Inputstream checks magic byte itself to identify bz2 (leave it in input) - try to suppy fitting mime for parsing bz2 content	2015-11-03 03:35:01 +01:00
luc	745e97a575	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-11-02 08:10:11 +01:00
reger	a60b1fb6c2	differentiate api call getLocalPort() from getConfigInt()	2015-10-31 23:09:03 +01:00
reger	11f3666660	increase use of pre.defined CATCHALL_QUERY string	2015-10-31 19:44:31 +01:00
reger	a58ee49307	Optimize internal imagequery focus on using content_type to select images (in favor of url file extension)	2015-10-31 19:18:46 +01:00
luc	fc3294382e	Updated javadocs for warning on target encoding format potential errors.	2015-10-30 16:19:05 +01:00
luc	aa70ff4ff6	Corrected images alpha channel rendering	2015-10-30 05:18:16 +01:00
reger	d223cf0ae4	adjust MediaWiki importer geo coordinate calculation - allow lat/long 0.xxx - south / west assignment include test class	2015-10-26 21:19:35 +01:00
reger	2b775d5be6	fix typo in WikiCode coordinate calculation	2015-10-25 19:38:42 +01:00
reger	bbe9df2bb3	fix MediawikiImporter for bz2 dump skip reading bz2 file magicbyte to identify bz2 format as inputstream reset would be required. Common compress reads and checks the magicbytes internally and throws ioexception if wrong, making preread obsolete.	2015-10-25 03:06:15 +01:00
reger	c6687dd560	fix a system.out to log.fine in bmpParser	2015-10-25 00:26:45 +02:00
reger	e53c6bbd51	fix init of peer flags (remove hiding of ssl flag)	2015-10-24 19:36:33 +02:00
Michael Peter Christen	ac034db8bc	Merge branch 'master' of https://github.com/luccioman/yacy_search_server # Conflicts: # htroot/js/highslide/highslide.js # source/net/yacy/document/ImageParser.java	2015-10-24 11:22:35 +08:00
reger	826f14f37f	fix unnececary set null of peer flags, causing reread remove obsolete version flags	2015-10-22 02:35:58 +02:00
luc	5902ce032e	Corrected NullPointerException case when ImageIO reader is not found for image format.	2015-10-19 14:11:26 +02:00
reger	c6495a5b62	add a log entry on parsing ajax crawling scheme snapshot (prev. commit `9252e36aeb`)	2015-10-18 06:19:12 +02:00
reger	9252e36aeb	implement ajax crawling scheme for ajax sites which adhere to the proposed use of hash-bangs to provide html content see freshly deprecated https://developers.google.com/webmasters/ajax-crawling/ Implementation improves parsing of the homepage (ajax page) which uses metatag "fragment" in header and parses supplied html snapshot instead of mostly empty ajax/scripted page. Implementation supports also hash-bang urls (url with anchor starting with ! like ...path#!hashfragment) but our crawler filters it (use of hash-bang is controversly discussed and proposal is deprecated, makes no sense to adjust the crawler, but as long as it is used by some sites the minor change/improvement in htmlparser is good for some time). Quick - how does it work - if metatag fragment with content "!" is found - htmlparser tries to get content of htmls snapshot (using a different url) - htmlparser returns 2 documents (original url and snapshot content - but using same original url) - after parsing result documents are joined (and stored to index containing content also from snapshot page... as the original ajax page contains typically no parseable html content)	2015-10-18 05:51:01 +02:00
Michael Peter Christen	d1ae999ef9	replaced HashMap with LinkedHashMap to preserve the object order	2015-10-16 23:30:51 +02:00
Michael Peter Christen	7d075a1d76	added log lines	2015-10-16 23:30:04 +02:00
Michael Peter Christen	092dac086e	Merge branch 'master' of https://github.com/luccioman/yacy_search_server	2015-10-16 23:22:30 +02:00
reger	7a64bebb86	init Recrawl job chunk size to max crawl loader during job start, to use some system preferences and allow injection of recrawl urls before queue is empty During recrawl the balancer hangs on the very last urls often on hosts with huge delay time, by allowing injection earlier progress is more balanced. Max number of injected crawl urls by recrawl job is 2 * max loader.	2015-10-16 03:05:39 +02:00
luc	d6522fa4a2	Integrated haraldk/TwelveMonkeys library to first add TIF image format support.	2015-10-15 10:06:51 +02:00
Michael Peter Christen	9244694e64	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-10-14 15:17:23 +02:00
Michael Peter Christen	151ccd50a9	fix for image size field values (must be multi-valued)	2015-10-14 15:16:16 +02:00
reger	c9937973e3	unescape MultiProtocolURL getAttributes() return values. use getAttributes() to get query parameters as clear text (w/o url encoding) use getSearchpartMap() to get in internal format (url encoded) fix for http://mantis.tokeek.de/view.php?id=606	2015-10-13 02:43:18 +02:00
reger	78e8c6f3e5	refactor special handling (static override) of SUPPORTED_EXTENSIONS/MIME_TYPES not used for genericImageParser	2015-10-11 01:23:52 +02:00
reger	d54c5d310a	add links with image extension not automatically to image links. With the wide spread use e.g. of Wikimedia the url file extension of links with image extension often point to html.	2015-10-10 23:49:58 +02:00
reger	851e8f6c8a	check jpeg file signature in genericImageParser to fail early without further object allocation if source is not a jpeg.	2015-10-05 01:58:31 +02:00
reger	fb75fea446	use recrawljob w/o sort results by date This is a workaround for existing index (not fully reindexed) since intro of schema with docvalues to prevent solr exception causing recrawljob to fail with org.apache.solr.core.SolrCore java.lang.IllegalStateException: unexpected docvalues type NONE for field 'load_date_dt' (expected=NUMERIC). Use UninvertingReader or index with docvalues.	2015-10-04 05:43:40 +02:00
reger	43c27aa550	upd to solr/lucene 5.3.1	2015-10-03 23:20:33 +02:00
reger	688f7b2a5c	allow/display svg images in image results previews svg is not supported by awt but by most browser. Image content is delivered as received (without size adjustment)	2015-10-02 01:48:48 +02:00
reger	d5330391de	remove some unused var allocation in parser	2015-10-01 23:11:58 +02:00
Michael Peter Christen	3d7dd9d3aa	follow-up to latest commit: also flush the search cache if all crawls had been terminated.	2015-10-01 13:21:28 +02:00
Michael Peter Christen	c737ff235d	in case that the include_string contains several entries including 1-char tokens and also more-than-1-char tokens, then remove the 1-char tokens to prevent that we are to strict. This will make it possible to be a bit more fuzzy in the search where it is appropriate.	2015-10-01 13:09:33 +02:00
Michael Peter Christen	8e555d79a3	add also 1-character tokens to the token list because that could be also searched for. A full-string search for a filename may fail if those 1-char tokens are omitted	2015-10-01 13:03:22 +02:00
reger	7c82cd4415	add a end condition to svgParser for wrong content (if parser choosen just by file extension)	2015-09-29 22:57:33 +02:00
reger	356d4d1301	remove rdfParser from init (current function identical with genericParser)	2015-09-26 17:30:34 +02:00
reger	c647d899e3	add svgParser to parse metadate from svg images Reads document level included title and description and skips the graphic content to save bandwidth. svg metadata element is not interpreted - remove rdfParser from init (current function identical with genericParser)	2015-09-26 17:27:33 +02:00
reger	bad34804fe	optimize parseInt for <img> tag attribute parsing Performance better as using Numberformat.parse or parseInt(substring())	2015-09-26 15:42:23 +02:00
Michael Peter Christen	6ebc2451a9	Merge pull request #14 from luccioman/master Translator refactoring : no more regular expression processing	2015-09-24 13:50:23 +02:00
reger	2f51baff4f	check for loading error (includs unsupported formats) to prevent blank thumbnail display in image search because of not handled source which don't load on click. Now the cross icon indicates the problem (inlcuding not supported format)	2015-09-24 01:58:19 +02:00
luc	5578886f6f	Merge branch 'master' of https://github.com/luccioman/yacy_search_server.git	2015-09-23 21:04:20 +02:00
luc	c38d6c1f37	Correction for mantis 535: inurl: parameter doesn't work on URLs with upper-case letters	2015-09-23 21:01:51 +02:00
reger	52e3eb4ce8	harmonize/correct assignment to Ymarkmeta.mime replace use of deprecated	2015-09-23 00:13:10 +02:00
Michael Peter Christen	87f358058e	Fix for index entries which have id's not computed as hash from the url. This makes it possible to operate with outside-computed url hashes in enterprise environments not using the build-in crawler from YaCy.	2015-09-22 11:56:17 +02:00
reger	3f2b8ab5e5	optionally include mime in p2p url exchange string if doctype decodes to ambiguous mime and default conversion is not equal to original	2015-09-22 00:12:31 +02:00
reger	a3195d78ae	add Portuguese month names to date recognition	2015-09-20 23:28:42 +02:00
reger	d2cc11ea8f	fix html parser taking <style> content as text. Noticed some result description contain css content from style tag. Added <style> to tag list to scrape it's content not as text + test case included	2015-09-19 05:30:55 +02:00
Michael Peter Christen	5f706797cb	patch for a bug inside of solr since solr 5.0 when using a boost function with a numeric date field: "unexpected docvalues type NUMERIC for field 'last_modified' (expected one of [SORTED, SORTED_SET]). Use UninvertingReader or index with docvalues." This is a well-known bug inside solr which prevents that now the 'sort by date' in the YaCy search interface can be used. Without this patch no results at all is displayed (since the exception prevents that). Now there is at least a result but it is not ordered properly.	2015-09-18 02:25:44 +02:00
reger	7889fc2389	Hack to prevent Solr issue on partial update on a document containing multivalued date field (regardless if these fields part of update). Switch partial update option off in postprocessing if schema contains *_dts (multivalued date field). see http://mantis.tokeek.de/view.php?id=601	2015-09-13 20:23:15 +02:00
reger	b4cbdea1e7	adapt SolrServerConnector.add to handle error on partial update input document. In case of error we deleted the original document and added the new doc to the index. This is not valid for partial update documents (which contain only a subset of the fields). Remove the "delete" error handling step.	2015-09-13 20:19:50 +02:00
reger	98ab655917	on reindex delete index document with invalid url if discovered	2015-09-12 23:06:13 +02:00
reger	1e8369e18b	use a parsed date in Document.toString	2015-09-12 22:00:40 +02:00
luccioman	199b2ce52d	Translator refactoring : to simplify locale files writing, process keys as simple string and no more as regular expressions. Updated all locale files to adapt to refectored Translator : removed useless escaped characters and did minor corrections. Performed minor syntax corrections on some html source files. Added an util to translate all html source files with all locales without launching full YaCy application. Corrected main arguments parsing on other translation utils.	2015-09-11 17:20:11 +02:00
luccioman	4dd9c0d5d9	Merge from main repository	2015-09-08 08:54:48 +02:00
reger	3428b6f13b	improve filtering by filetype navigator. The used url-filter for filetype doesn't require ".ext" resulting in too many matches, add a sort-out filter for RWI results.	2015-09-07 02:36:22 +02:00
reger	e37a4f0b3d	prevent metadata records in index w/o valid url by throwing MalformedURL exception on URIMetadataNode creation	2015-09-06 22:19:05 +02:00
reger	41c4eade51	extract modification date from vCard (vcfParser)	2015-09-06 04:28:27 +02:00
reger	8768896975	extract lastmodified from openoffice doc set lastmod date in office document parsers	2015-09-06 00:04:54 +02:00
Michael Peter Christen	c40c302748	when many crawl queues are generated, this NPE can occur; probably caused as concurrency issue: W 2015/09/05 14:09:10 ConcurrentLog java.lang.NullPointerException java.lang.NullPointerException at java.util.TreeMap.rotateRight(TreeMap.java:2239) at java.util.TreeMap.fixAfterInsertion(TreeMap.java:2271) at java.util.TreeMap.put(TreeMap.java:582) at net.yacy.kelondro.table.Table.<init>(Table.java:235) at net.yacy.crawler.HostQueue.openStack(HostQueue.java:229) at net.yacy.crawler.HostQueue.getStack(HostQueue.java:204) at net.yacy.crawler.HostQueue.push(HostQueue.java:397) at net.yacy.crawler.HostBalancer.push(HostBalancer.java:237) at net.yacy.crawler.data.NoticedURL.push(NoticedURL.java:184) at net.yacy.crawler.CrawlStacker.stackCrawl(CrawlStacker.java:355) at net.yacy.crawler.CrawlStacker.job(CrawlStacker.java:134) at sun.reflect.GeneratedMethodAccessor6.invoke(Unknown Source) at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.lang.reflect.Method.invoke(Method.java:483) at net.yacy.kelondro.workflow.InstantBlockingThread.job(InstantBlockingThread.java:101) at net.yacy.kelondro.workflow.AbstractBlockingThread.run(AbstractBlockingThread.java:82) at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) at java.util.concurrent.FutureTask.run(FutureTask.java:266) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) at java.lang.Thread.run(Thread.java:745)	2015-09-05 14:12:17 +02:00
reger	367fe388b9	fix exception throw after sendError in DefaultServlet - reduce debug exception logs in crawler	2015-09-05 01:57:30 +02:00
luccioman	9752bd5f88	Added utils to help translation without launching full YaCy application : - translate all source files with a locale - list all non translated files with a locale	2015-09-04 13:44:44 +02:00
luccioman	2f0f0180e2	Added a function to list files recursively.	2015-09-04 13:42:57 +02:00
luccioman	7e4c1d2282	Translator refactoring : - deleted useless new StringBuilder allocation - use of a new reusable FileNameFilter - added javadoc	2015-09-04 13:42:10 +02:00
reger	802ccaead6	fix init of error cache, use latest faildates => load_date_dt	2015-09-02 02:36:31 +02:00
reger	dba7f15073	apply same size constrain on result image from doc as for linked images see `19f1308bf0`	2015-09-01 23:22:48 +02:00
reger	4cf875336c	complete TODO: getFileExtension handle dot in query part + testcase	2015-08-31 23:28:03 +02:00
sixcooler	87e4abe393	fight the fieldcache by usind DocValues: in Solr-5.x the fieldcache has moved and was not cleared anymore. This results in an huge fieldcache. (http://lucene.apache.org/#highlights-of-the-lucene-release-include https://issues.apache.org/jira/browse/LUCENE-5666) Here I try to use DovValues where it is possible. For this I used the Api-Scheme as new basis für the Solr-Schema. This needs at least a complete optimization of the Solr-Index to get a smaller FieldCache. Everything that is indexed with these setting will not use the Fieldcache at all.	2015-08-31 20:24:41 +02:00
reger	eaf0e8ff2c	start recording/indexing pixel size for image document as for linked images	2015-08-31 01:58:36 +02:00
reger	c33229fc0c	check mime prior to ext for metadata modification for images	2015-08-30 23:02:19 +02:00
reger	19f1308bf0	enforce th result images limit to > 16x16px for linked images http://mantis.tokeek.de/view.php?id=594	2015-08-30 02:19:52 +02:00
reger	0e4ba0360b	fix NPE on .yacyh result url of disconnected peer (cleanup yacyshare remaining)	2015-08-25 23:26:17 +02:00
reger	7ed812a2bf	log missing seed.port in favour of exception to prevent repeating throws	2015-08-25 02:19:00 +02:00
reger	206883f80d	fix: Preserve protocol in url proxy to connect to http/https. Display warning if https target is viewed over http	2015-08-25 01:16:41 +02:00
reger	f7b0b3b7b3	avoid runtime exception by earlier testing for seed.ip=null	2015-08-23 23:01:20 +02:00
Michael Peter Christen	906b5fd742	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-08-11 00:42:46 +02:00
Michael Peter Christen	8f90767889	fix for filesystem crawl	2015-08-11 00:42:26 +02:00
sixcooler	a3dd4be749	added / corrected charste to be 1.7 compatible. @Orbiter: please check is this is ok for you	2015-08-10 20:53:20 +02:00
Michael Peter Christen	8028410ab7	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-08-10 14:27:53 +02:00
Michael Peter Christen	df3314ac1a	added a new facet type based on a probabilistic classifier using bayesian filters. This can be used to classify documents during indexing-time using a pre-definied bayesian filter. New wordings: - a context is a class where different categories are possible. The context name is equal to a facet name. - a category is a facet type within a facet navigation. Each context must have several categories, at least one custom name (things you want to discover) and one with the exact name "negative". To use this, you must do: - for each context, you must create a directory within DATA/CLASSIFICATION with the name of the context (the facet name) - within each context directory, you must create text files with one document each per line for every categroy. One of these categories MUST have the name 'negative.txt'. Then, each new document is classified to match within one of the given categories for each context.	2015-08-10 14:27:44 +02:00
reger	1409cabe8b	exclude more default search fields from text copy to text_t for metadata index documents	2015-08-09 21:01:30 +02:00
reger	e2e73258ca	remove obsolete interface SearchAccumulator and unused SRURSSConnector Thread inheritance	2015-08-08 18:35:49 +02:00
Michael Peter Christen	dbbad23e12	removed warnings	2015-08-03 05:37:34 +02:00
Michael Peter Christen	500cfa9457	enhanced logging	2015-08-03 05:17:22 +02:00
Michael Peter Christen	c14bc8d9b7	revert of fq transformation (recent fix)	2015-08-03 05:15:34 +02:00
Michael Peter Christen	203df5a750	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-08-03 05:02:26 +02:00
reger	fa08ca207e	! finish running crawls before applying ! Allow crawl urls up to 2048 character fix for http://mantis.tokeek.de/view.php?id=575	2015-08-03 00:49:24 +02:00
reger	ee77f24e52	use some more declared HeaderFramework constants	2015-08-02 22:56:14 +02:00
Michael Peter Christen	11a848da5a	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-08-02 14:53:36 +02:00
Michael Peter Christen	b94bd7f20a	a collection of search query enhancements: - fixed superfluous space in query field list - fixed filter query logic - removed look-ahead query which caused that each new search page submitted two solr queries - fixed random solr result orders in case that the solr score was equal: this was then re-ordered by YaCy using the document hash which came from the solr object and that appeared to be random. Now the hash of the url is used and the score is additionally modified by the url length to prevent that this particular case appears at all.	2015-08-02 14:52:41 +02:00
reger	dbe2594c38	replace deprecated myPublicLocalIP() in AbstractRemoteHandler	2015-08-02 00:53:49 +02:00
reger	6d3534e725	remove unused Transmission hit counter	2015-08-02 00:20:14 +02:00
reger	cb67eb7baf	use more absolute path for config file opening as suggested in pull request 5 (https://github.com/yacy/yacy_search_server/pull/5)	2015-08-01 23:54:26 +02:00
Michael Peter Christen	1ccbf739b1	added bayes filter from Philipp Nolte, originally taken from https://github.com/ptnplanet/Java-Naive-Bayes-Classifier and modified inside the loklak.org project. After optimization in loklak it was inserted into the net.yacy.cora.bayes package. It shall be used to create custom search navigation filters. The original copyright notice was copied from the README.md from https://github.com/ptnplanet/Java-Naive-Bayes-Classifier/blob/master/README.md The original package domain was de.daslaboratorium.machinelearning.classifier	2015-07-30 14:10:31 +02:00
Michael Peter Christen	1bced1ae60	using latest enhanced (un/)gzip methods from loklak for yacy	2015-07-30 13:39:10 +02:00
Michael Peter Christen	3e6657288d	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-07-30 03:39:11 +02:00
Michael Peter Christen	de8cfbe1d7	added export option to export the fulltext of the search index text only	2015-07-30 03:21:40 +02:00
reger	2fb6ebe88a	move java environment parameter setting disabling SNI (Server Name Indicator) support for https connections from code to startup script allowing admin to ~easy/transparent alter the YaCy default FALSE setting. Background: some user report problem with connecting/crawling some sites via https which require SNI support (by default switched off in YaCy). On the other hand systems not demanding SNI support are sometimes not properly configured and due to a bug/feature in java 1.7 connection is aborted. The later is more often the case, so the default is still fine. With the java start parameter expert user can no alter the startparameter to -Djsse.enableSNIExtension=true (java default) if they crawl more hosts requiring SNI support. The alternative to let YaCy try both during https handshake (deep inside the httpclient) is not pursut at this time.	2015-07-29 23:30:05 +02:00
Michael Peter Christen	fbeae20b3a	try a healing of the cache if the index file is corrupted	2015-07-27 15:16:08 +02:00
Michael Peter Christen	03ea723889	added log lines for query performance profiling	2015-07-27 15:03:13 +02:00
Michael Peter Christen	0e87a99ab8	more fixes for special windows paths	2015-07-10 17:34:29 +02:00
Michael Peter Christen	e5b6424eed	patch for bad windows file paths	2015-07-10 17:14:14 +02:00
Michael Peter Christen	0aa6fcf259	remove old vocabularies and synonyms before adding new	2015-07-10 16:47:19 +02:00
Michael Peter Christen	289018b559	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-07-08 17:37:03 +02:00
Michael Peter Christen	7b412e8c07	added msg (text emails) format; should be handled by html parser.	2015-07-08 17:36:37 +02:00
reger	f91298d3b6	fix one implicit Integer/Long type conversion -> causes Java 1.8 compile error	2015-07-08 03:02:10 +02:00
reger	821262a179	add CommonPattern for multiple spaces to eliminate empty split words on following spaces	2015-07-04 22:49:01 +02:00
Ryszard Goń	59096935d0	Use language-detection library for increased accuracy	2015-07-02 18:41:13 +02:00
Michael Peter Christen	90f75c8c3d	added enrichment of synonyms and vocabularies for imported documents during surrogate reading: those attributes from the dump are removed during the import process and replaced by new detected attributes according to the setting of the YaCy peer. This may cause that all such attributes are removed if the importing peer has no synonyms and/or no vocabularies defined.	2015-07-02 00:23:50 +02:00
Michael Peter Christen	7829480b82	refactoring: separated condenser and tokenizer	2015-07-01 18:28:18 +02:00
Michael Peter Christen	593de05922	enhanced surrogate import process speed (dramatically!)	2015-06-29 12:28:34 +02:00
Michael Peter Christen	3c4c69adea	fix for - bad regex computation for crawl start from file (limitation on domain did not work) - servlet error when starting crawl from a large list of urls	2015-06-29 02:02:01 +02:00
Michael Peter Christen	1fec7fb3c1	suppress access to solr when doing search suggestions in case that the index has more than two million documents. This protects the index from beeing flooded with search requests that cannot be resolved before the real search query has to be computet.	2015-06-24 13:02:12 +02:00
Michael Peter Christen	694b22f165	migration to Solr 5.2: huge benefits - this is a lot faster! This is a very complex migration: many classes had been renamed or removed, dependencies changed and the solr index type is now aligned to be a solr cloud repository. Together with the Solr 5.2 library update, one other dependent library had been updated as well: httpclient 4.4->4.4.1 Older indexes are migrated from 4_10 to 5_2. However, the new index structure is more efficient and we recommend to re-index everything. Please use the index export before you do the update to a large surrogate xml file. After the update, start with an empty index and then initialize this with your dump.	2015-06-24 01:55:51 +02:00
sixcooler	e427efbe54	Next Try for a fix for upload-connection staying in blocked state. This was caused by reading via GZIP from close-wait connection an caused high cpu- and system-loads. Instat of implementing handling of the RedListener now I found a timelimeted 'get' "realy" solving this problem.	2015-06-14 22:56:26 +02:00
reger	0fab445b19	Resourceobserver log warning - deleting releases files - only on actual deletes instead of entering routine	2015-06-10 02:35:37 +02:00
sixcooler	ef6a64b2a4	Fix for upload-connection staying in blocked state. This was caused by reading via GZIP from close-wait connection an caused high cpu- and system-loads. Solved by implementing handling of the RedListener.	2015-06-09 21:26:10 +02:00
reger	c973f94936	add log entry on release file delete by ResourceObserver	2015-06-08 03:17:12 +02:00
reger	121972752c	implement deleteOldDownloads in RexourceObserver on low diskspace - direct assign sb.observer (skip redundant InitThread)	2015-06-08 02:52:13 +02:00
Michael Peter Christen	9c12555be5	added link to Snapshots in search results if the snapshot exists and option is set in ConfigSearchPage_p (this is a stub: we also need a visualization of pdf files!)	2015-06-07 20:37:37 +02:00
reger	72f6a0b0b2	enhance recrawl job - allow to modify the query to select documents to process (after job has started) - allow to include failed urls (httpstatus <> 200)	2015-06-06 18:45:39 +02:00
reger	7478338a40	remove augmented parsing activation from frontend experimental implementation not used and based on error prone experimental rdfaparser	2015-06-05 00:51:00 +02:00
reger	11aa2edfe1	remove RDFa parser activation from frontend reason: experimental implementatin of RDFa parser not executed (limited to special urls) but may cause error on normal html parsing due to a inputstream.reset	2015-06-05 00:15:16 +02:00
reger	49b79987c9	remove obsolete searchfl work table was used to register urls with not complete words in snippet but is never accessed	2015-06-04 22:44:01 +02:00
Michael Peter Christen	d0aff91f23	fix for index import	2015-06-01 01:56:09 +02:00
Michael Peter Christen	34de1e8cbc	gzip compression will perform more efficient and with better compression level	2015-06-01 01:24:33 +02:00
Michael Peter Christen	98be59ce9c	full solr xml exports will now be automatically compressed during export. That makes it possible to export a solr xml dump even if disc space is low.	2015-05-30 19:02:54 +02:00
Michael Peter Christen	a1a8edfc0a	wrap HeaReader close() in a catch Throwable block to prevent that an excpetion during close blocks the whole shotdown process	2015-05-30 17:54:02 +02:00
Michael Peter Christen	b43811d38c	added surrogate import process for exported solr dumps. Just throw your solr dump file into DATA/SURROGATES/in/ and it will be imported!	2015-05-30 13:19:59 +02:00
Michael Peter Christen	b77537294d	prevent disc usage when showing tray animation	2015-05-30 06:57:15 +02:00
Michael Peter Christen	eec78e1b0c	added intensity option to graphics	2015-05-30 06:31:08 +02:00
Michael Peter Christen	a5007f345e	re-licensing some of my old visualization classes under LGPL 2.1	2015-05-30 06:12:08 +02:00
Michael Peter Christen	c99a665593	adding a 3-pixel font generator made some time ago..	2015-05-30 06:01:52 +02:00
Michael Peter Christen	c7576d6028	added a full solr export to the IndexControlURLs_p.html servlet. The export function is also now the default export option. The export file format for a full solr export is very similar to a solr search result xml, only the <lst name="responseHeader"> tag is missing. The exported xml has a special line termination feature: all documents will be exported into a single line without any CR in between. That means that every document is completely inside a single line. While this is not readable at all for humans, it is very useful for linux line processing scripts, like grep. Using grep it will be easy to select single documents which match for a given pattern. Such dumps shall be importable with the DATA/SURROGATE/in import function, but that import is not yet adopted to the new file format.	2015-05-29 15:05:52 +02:00
Michael Peter Christen	197f7449e5	All entities of crawl profiles are now editable in the crawl profile editor.	2015-05-28 16:07:40 +02:00
reger	1d8e1e4bac	- Image search expand box, adjust javascript hs padtominsize parameter, to make sure expand box doesn't shrink on small images - asure ImageResult.imagetext has value for the link text (use filename if no alt text given)	2015-05-27 02:31:13 +02:00
reger	8b35656007	remove hard throw exception in makeResultEntry remove not used "share." peername.yacy url rewrite	2015-05-26 23:57:06 +02:00
reger	af57fbefad	use available mime (instead null) on imageresult from metadatanode	2015-05-26 23:54:04 +02:00
reger	dd7782bac0	revert deletion of BinSearch (accident)	2015-05-26 04:26:26 +02:00
reger	000dde9511	Eleminate duplication of values for search ResultEntry by instatiation from URIMetadataNode, by eleminating differentiation of ResultEntry/URIMetadataNode. - moved remaining ResultEntry functionallity to URIMetadataNode - for 1:1 functionallity added a function makeResultEntry() - removed ResultEntry - refactored related code Main difference is after makeResultEntry the text_t content is removed and alternative title/url strings for display are calculated. Main difference left is, that	2015-05-26 04:15:00 +02:00
reger	29c4aa3991	fix compiler notification of missing serialID from last commit	2015-05-25 21:51:32 +02:00
reger	3d53da8236	refactor ResultEntry to be based on MetadataNode/SolrDocument to share/reuse common access routines	2015-05-25 21:28:48 +02:00
reger	d882991bc5	Implement sharing of ioDispatcher for term & citation index as proposed in ioDispatcher description	2015-05-25 19:46:26 +02:00
reger	370ba9da71	On imageSearch prefere mime to sort out none-image documents Generalize the hack to prevent urls with just a img extension beeing returned improving http://mantis.tokeek.de/view.php?id=528	2015-05-24 21:48:58 +02:00
reger	cd31633369	improve MultiprotocolURL.getFileExtension() prevent string OOB while querypart contains a dot (return just "") see log snippet in http://mantis.tokeek.de/view.php?id=533	2015-05-24 19:38:04 +02:00
reger	c60ccdfbcf	Increase IODspatcher dumpQueue size to 2 to reduce risk of concurrent emergency dump, skip concurrent emergency merge dealing with/see http://mantis.tokeek.de/view.php?id=566	2015-05-24 18:03:27 +02:00
reger	8a9622c31c	fix string OoB on getImagelinks with long alttext in description calculation	2015-05-24 01:59:40 +02:00
reger	3e742d1e34	Init remote crawler on demand If remote crawl option is not activated, skip init of remoteCrawlJob to save the resources of queue and ideling thread. Deploy of the remoteCrawlJob deferred on activation of the option.	2015-05-23 02:06:39 +02:00
reger	13f013f64a	Limit extra sleep of BusyThread on LowMemCycle	2015-05-17 06:21:12 +02:00
reger	cd7c0e0aae	detail optimization of RecrawlThread	2015-05-17 00:13:00 +02:00
reger	ace71a8877	Initial (experimental) implementation of index update/re-crawl job added to IndexReIndexMonitor_p.html Selects existing documents from index and feeds it to the crawler. currently only the field fresh_date_dt is used determine documents for recrawl (fresh_date_dt:[* TO NOW-1DAY] Documents are added in small chunks (200) to the crawler, only if no other crawl is running.	2015-05-16 01:23:08 +02:00
reger	141cd80456	correct log msg text	2015-05-16 00:01:54 +02:00
reger	f3ce99bfb8	fix extract of inboundlinks_protocol_sxt url counter maybe > 999	2015-05-14 00:03:09 +02:00
reger	2bc9cb5828	fix early return in addToCrawler check / handle all supplied urls after error url	2015-05-13 21:58:43 +02:00
Michael Peter Christen	f5f88272e4	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-05-12 12:06:42 +02:00
Michael Peter Christen	5c67c4d460	fix for latest commit, see `f810915717 (commitcomment-11145880)`	2015-05-12 12:06:21 +02:00
reger	c37dda8849	fix NPE on MultiProtocolURL on url with parameter value and '=' in getAttribute - added test case for it	2015-05-12 01:09:10 +02:00
Michael Peter Christen	f810915717	added crawl start from a clone with very, very large url: they are now encoded as post submit form inside a javascript creation function.	2015-05-11 16:30:41 +02:00
Michael Peter Christen	51de86c992	disabled debug thread dumps	2015-05-11 14:46:09 +02:00
Michael Peter Christen	d524a9d77c	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-05-11 14:42:40 +02:00
Michael Peter Christen	0710648c31	enable api calls with very long urls	2015-05-11 14:42:21 +02:00
reger	31346e873b	upd library reference of missing jsch-0.1.21 in seeduploadscp.xml upd to jsch-0.1.52.jar	2015-05-11 01:35:12 +02:00
reger	609c52e987	refactor getBookmark to consistenly check existance by != null (w/o throwing exception on not found)	2015-05-11 00:37:04 +02:00
reger	1481a8ab56	add opensearch rss results to dht collection (due to text = snippet) which is used to differentiate meta from full data - make sure check for dht is not dependant on number of collection entries	2015-05-10 18:52:33 +02:00
reger	f134aa7f7f	persist bookmark timestamp on setTimeStamp()	2015-05-10 15:29:23 +02:00
reger	752eec6697	fix NPE in addToIndex when used outside searchEvent	2015-05-10 05:18:23 +02:00
Michael Peter Christen	fbf85a1561	added temporary debug output in http client	2015-05-08 15:31:01 +02:00
Michael Peter Christen	ff29b0e503	added option to re-index exported xml snapshot dumps to HTCACHE/snapshots by just placing them in the SURROGATES/in path	2015-05-08 15:30:26 +02:00
Michael Peter Christen	6f4fe4b175	revert of `8a7c68e4c7` keeping surrogates after processing is essential for some users. If the space they are taking is too high, please set up an automatic deletion process (like a cronjob).	2015-05-08 14:01:30 +02:00
Michael Peter Christen	97930a6aad	added must-not-match filter to snapshot generation. also: fixed some bugs	2015-05-08 13:46:27 +02:00
Michael Peter Christen	9d8f426890	adding a try-catch to link graph processing to prevent that a single malformed url interrupts the storage process	2015-05-08 10:38:33 +02:00
reger	8a5b8f8789	on bookmaring of search result, remember orig. query in separate bookmark property (instead of using the description field) - adjust display and autosearch - don't overwrite existing bookmark but combine info	2015-05-03 02:31:50 +02:00
reger	7224209486	break out of NormalizeDistributor loop on timeout	2015-05-02 02:36:18 +02:00
reger	47e61f8325	fix typo in image filter query (extra bracket)	2015-04-28 03:12:14 +02:00
reger	4b4ab6799f	fix String out of range in Collection Nav see http://mantis.tokeek.de/view.php?id=573	2015-04-27 22:38:40 +02:00
reger	572cfe8fd4	improve character encoding for urlproxy servlet for none utf-8 pages	2015-04-26 17:42:39 +02:00
reger	6bc8a9b11e	make Quality of Service Servlet available to prioritize requests from local host This assigns priorities to incoming requests. Higher priority numbers are served before lower. (disabled by default in defaults/web.xml, uncomment or copy entry to DATA/Settings/web.xml)	2015-04-26 04:29:32 +02:00
Ryszard Goń	ca1a70aec8	fix for Accept '?' URLs column in Crawl Profile List	2015-04-19 15:55:49 +02:00
reger	5408448a56	skip redundant add. of keywords to text search uses keywords as default search field	2015-04-17 02:14:13 +02:00
reger	296e97c78e	put https port in peers dna as we flag if a peer is accesible via https, we need to know the port if we want to use is (e.g. for interYaCy communication) start to provide / tansport the port by recording it in peers dna. - add https link on the Network.html lock symbol	2015-04-16 02:36:12 +02:00
Michael Peter Christen	fed26f33a8	enhanced timezone managament for indexed data: to support the new time parser and search functions in YaCy a high precision detection of date and time on the day is necessary. That requires that the time zone of the document content and the time zone of the user, doing a search, is detected. The time zone of the search request is done automatically using the browsers time zone offset which is delivered to the search request automatically and invisible to the user. The time zone for the content of web pages cannot be detected automatically and must be an attribute of crawl starts. The advanced crawl start now provides an input field to set the time zone in minutes as an offset number. All parsers must get a time zone offset passed, so this required the change of the parser java api. A lot of other changes had been made which corrects the wrong handling of dates in YaCy which was to add a correction based on the time zone of the server. Now no correction is added and all dates in YaCy are UTC/GMT time zone, a normalized time zone for all peers.	2015-04-15 13:17:23 +02:00
Michael Peter Christen	b060ba900d	added parsing of contentprop attribute in html tags for content='startDate' and content='endDate'. The value of these field is now written to new solr fields startDates_dts and endDates_dts.	2015-04-13 16:20:00 +02:00
Michael Peter Christen	4cb4f67f38	added parsing of dd, dt and article html fields. The parsed result is written to special solr fields which are deactivated by default.	2015-04-12 22:02:45 +02:00
reger	1395f10e95	fix typecast for css links	2015-04-12 01:11:47 +02:00
Michael Peter Christen	3288489fd2	more logging during start-up	2015-04-11 13:00:32 +02:00
Michael Peter Christen	abaaaef5f1	fix for filter queries	2015-04-11 12:30:29 +02:00
Michael Peter Christen	4d00175157	<experimental> added parsing of <article> html element. Whenever such an element occurs, the complete content of all article elements replaces the parsed <content> part of documents.	2015-04-10 16:16:20 +02:00
Michael Peter Christen	1df6492019	enhanced suggestions	2015-04-10 15:59:18 +02:00
Michael Peter Christen	ae02c92fd0	logging fix	2015-04-09 14:21:23 +02:00
Michael Peter Christen	5651713134	better debugging of fq	2015-04-07 17:02:02 +02:00
Michael Peter Christen	f5a032f293	split query into filter query and text query to get better ranking results and faster results	2015-04-07 16:10:13 +02:00
Michael Peter Christen	2e88028c1a	when selecting collections in navigation, do show the un-selected collections in search result. When selecting one of them in another search, switch off the previously selected collection. This actually turns the collection navigation modifier into a radio-button like behaviour	2015-04-07 13:13:58 +02:00
Michael Peter Christen	1de9b21c65	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-04-07 12:40:43 +02:00
reger	5f4cd8d6f5	replace deprecated getIP with getIPs in AbstractRemoteHandler	2015-04-07 00:10:42 +02:00
Michael Peter Christen	fa7edc9f7a	refactoring of filter queries (several queries instead only one)	2015-04-02 13:27:47 +02:00
Michael Peter Christen	40389987ec	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-04-01 18:18:05 +02:00
Michael Peter Christen	f9ba50379d	added an expansion option to search facets on result page: - if less or equal of 8 facet options are present, they are shown by default - if more facet options are present, they are hidden To view or hide all facets, just click on the facet header bar	2015-04-01 18:17:52 +02:00
reger	1f0f77bb77	make location facet return results for location nav facet of field coordinate_p does not return results, now using coordinate_p_0_coordinate as alternative to get facet counts. As the actual facet value is not used this should not harm any analysis (even if facet is a incomplete location). If facet value is used in future likely *_geohash field could be introduced (for facet and other ... as transport value)	2015-04-01 01:57:56 +02:00
reger	b1ec0644e5	fix NPE in location search on missing/empty PubDate in underlaying rss data	2015-03-31 02:20:13 +02:00
reger	c1dcc8c456	fix display and limit of max server connections after startup (on restart value returned to default=50) This has no effect on Jetty but the limit is still respected.	2015-03-29 07:12:23 +02:00
reger	839b962c20	correct percent encoding for '%' char	2015-03-28 03:05:21 +01:00
Michael Peter Christen	9bf0d7ecb9	added a new collection type 'dht' to all documents from the peer-to-peer interface to distinguish rich and poor document data. This also reverts some changes from commit `796770e070` because the firstSeen database is the wrong method to distinguish these types of data	2015-03-24 12:32:39 +01:00
reger	796770e070	prevent overwrite of crawled or received full documents by (newer) metadata To protect rich index data (full resource) from overwriting by metadata gathered during remote search, the newly introduced "firstSeen" index is used to differentiate between full-resource-doc and metadata, as a "firstSeen" entry is only added on store's of full-resource-docs (during crawl or remote search).	2015-03-23 03:57:47 +01:00
Michael Peter Christen	ee2490ab98	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-03-19 10:42:57 +01:00
reger	431311df42	fix get fresh_date_dt to allow returned value to be date in future	2015-03-18 22:04:03 +01:00
otter	74c7e8b686	Fixes hanging FlushThread (see http://forum.yacy-websuche.de/viewtopic.php?f=5&t=5447) by replacing put() method by the more robust add() to add a merge job to the queue.	2015-03-18 21:57:41 +01:00
reger	f63fff9008	fix snippet containig number with comma as desmo point http://mantis.tokeek.de/view.php?id=344 to keep it as one word (by altering the split regex) - added sniipet test case with number - regex for word split to match multiple splitcars	2015-03-16 02:03:40 +01:00
reger	b241264632	fix error on *abc query input http://mantis.tokeek.de/view.php?id=486	2015-03-15 22:31:47 +01:00
reger	2ef8ffdb60	apply UTF-8 encoding copied from escape()	2015-03-15 06:02:45 +01:00
reger	7120ea42f1	fix for path with char code > 255 (causing index out of bound exception) + test cas for it	2015-03-15 03:37:32 +01:00
reger	1d81bd0687	fix url encoding for path see http://mantis.tokeek.de/view.php?id=559 So far we used same escape procedure for all parts of the url (which includes x-www-form-urlencoded for all url components) Added capability to use different encoding rules for the different url components (through specific bitset for each component). (this is inspired by org.apache.http.client and java.net.uri implementation). - Added test case for http://mantis.tokeek.de/view.php?id=559	2015-03-15 00:46:07 +01:00
reger	62087fb8b2	fix MultiProtocolURL mailto protocol detection	2015-03-13 02:02:53 +01:00
reger	2e8c24e02a	fix link to DeReWo download file	2015-03-11 20:02:23 +01:00
reger	706f75ddc2	try to fix hang on index blob merge on shutdown http://mantis.tokeek.de/view.php?id=505 It happens but not able to reproduce. This change makes sure terminate signal is catched at end of currently running merge jobs	2015-03-11 19:36:23 +01:00
reger	f94e34058c	fix url (path) %-decoding http://mantis.tokeek.de/view.php?id=519 - add test case for this	2015-03-11 01:05:14 +01:00
reger	7e09bff4a1	exclude default search fields from text copy to text_t for metadata index documents (reduce text redundance)	2015-03-08 21:49:23 +01:00
reger	86073a5ba3	For remote crawlReceipt add document abstract/description enhance the returned metadata returned to the originator by description_txt to improve fulltext search result hits.	2015-03-08 02:34:48 +01:00
reger	8af70950d9	harmonize snippet computation to considere description_txt always (solr hl & internal). For now just added desc to text list for computation, could be further equalized with hl computation.	2015-03-05 02:22:05 +01:00
Michael Peter Christen	fd4e2c809a	Show dates in the content of a document in the search result: - if an eventDate is given in the search result, replace the document date with the event date and prefix it with the string "on ". - the document date is omitted if a date from the cent is shown Added also the date as fields in the json and rss result sets.	2015-03-02 18:00:20 +01:00
Michael Peter Christen	893889bc7b	added special terms for on: - Date modifier: tomorrow, today; i.e.: search for: "Berlin on:tomorrow" to find events happening tomorrow in Berlin	2015-03-02 13:10:05 +01:00
Michael Peter Christen	710a0efa1b	generalized time period computations	2015-03-02 12:55:31 +01:00
Michael Peter Christen	d9d3111d10	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-03-02 04:31:05 +01:00
Michael Peter Christen	535f1ebe3b	added a new way of content browsing in search results: - date navigation The date is taken from the CONTENT of the documents / web pages, NOT from a date submitted in the context of metadata (i.e. http header or html head form). This makes it possible to search for documents in the future, i.e. when documents contain event descriptions for future events. The date is written to an index field which is now enabled by default. All documents are scanned for contained date mentions. To visualize the dates for a specific search results, a histogram showing the number of documents for each day is displayed. To render these histograms the morris.js library is used. Morris.js requires also raphael.js which is now also integrated in YaCy. The histogram is now also displayed in the index browser by default. To select a specific range from a search result, the following modifiers had been introduced: from:<date> to:<date> These modifiers can be used separately (i.e. only 'from' or only 'to') to describe an open interval or combined to have a closed interval. Both dates are inclusive. To select a specific single date only, use the 'to:' - modifier. The histogram shows blue and green lines; the green lines denot weekend days (saturday and sunday). Clicking on bars in the histogram has the following reaction: 1st click: add a from:<date> modifier for the date of the bar 2nd click: add a to:<date> modifier for the date of the bar 3rd click: remove from and date modifier and set a on:<date> for the bar When the on:<date> modifier is used, the histogram shows an unlimited time period. This makes it possible to click again (4th click) which is then interpreted as a 1st click again (sets a from modifier). The display feature is NOT switched on by default; to switch it on use the /ConfigSearchPage_p.html servlet.	2015-03-02 04:30:10 +01:00
reger	d7259419f3	postpone raw snippet html encoding upon use instead of during init of snippet adressing http://mantis.tokeek.de/view.php?id=551	2015-02-28 19:02:18 +01:00
reger	de56d934b2	apply query parameter getQueryFields() to GSA servlet	2015-02-27 00:53:20 +01:00
reger	2d2299f484	fix mimetype of rss items in rss parser - remove self reference as anchor for items	2015-02-25 01:58:42 +01:00
Michael Peter Christen	b432049d59	enhanced date parsing time	2015-02-25 01:05:46 +01:00
reger	9b0de2de64	introduce getQueryFields to return default query fields (queryparamter QF) calculated from boostfields config, making sure title, description, keywords and content is always searched. - apply change to solrServlet makes sure every remote query uses at least all locally defined boost fields for search - apply to local solr search - simplify select query by using QF defaults	2015-02-23 23:12:07 +01:00
reger	a0f04db9ea	add extracted description/subject to pptParser	2015-02-22 05:31:56 +01:00
reger	8ec1db76ee	url unescape add check for inconsistent utf8 multibyte parsing If the url contains special chars (like umlaute äöü) it's interpreted as multybyte char and actually not converted at all (removed). Added a check if the multibyte convesion is not complete, just add the char as is. This fixes http://mantis.tokeek.de/view.php?id=200	2015-02-20 02:21:04 +01:00
reger	4b97ddb9ec	stop sending crawl receipts if receiver got offline	2015-02-17 03:16:10 +01:00
reger	7e35518787	add extracted description/subject to docParser	2015-02-16 00:50:16 +01:00
reger	f0a5188e11	replace depreciated HTTPClient setStaleConnectionCheckEnabled with setValidateAfterInactivity()	2015-02-15 23:09:01 +01:00
reger	7b569d2dbe	replace depriciated HTTPClient ALLOW_ALL_HOSTNAME_VERIFIER with NoopHostnameVerifier()	2015-02-15 21:34:01 +01:00
reger	fba34e12ef	fix formatting issue if snippet contains html code replacement for reverted commit `61f42a7928`	2015-02-15 20:39:20 +01:00
reger	e48720a58c	fix NPE in snippet computation	2015-02-15 05:30:14 +01:00
reger	eda0aeaf26	allow/recognize host in file: protocol crawl target This is useful in intranet indexing while crawling a intranet file server accessed via hostname while e.g. under Windows mapped to different drive letters on individual clients. Here you can crawl e.g. file://fileserver/documents having a valid uri in that intranet environment (while e.g. P:/documents might be client dependant).	2015-02-11 23:26:39 +01:00
reger	df83fcc4fc	disable optimistic GC assumption in StandardMemoryStrategy After several tests found that eom is not prevented. Major reason in testing was assumption future GC will free avg of last 5 GC. Disabeling this check improved eom exceptions. Added simplest testcase used for verification	2015-02-11 01:42:01 +01:00
Michael Peter Christen	8ff76f8682	the cleanup process experienced a 100% CPU load situation and the loop did not terminate: Occurrences: 100 at java.util.HashMap$KeyIterator.next(HashMap.java:956) at net.yacy.cora.protocol.ConnectionInfo.cleanup(ConnectionInfo.java:300) at net.yacy.cora.protocol.ConnectionInfo.cleanUp(ConnectionInfo.java:293) at net.yacy.search.Switchboard.cleanupJob(Switchboard.java:2212) at sun.reflect.GeneratedMethodAccessor12.invoke(Unknown Source) at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.lang.reflect.Method.invoke(Method.java:606) at net.yacy.kelondro.workflow.InstantBusyThread.job(InstantBusyThread.java:105) at net.yacy.kelondro.workflow.AbstractBusyThread.run(AbstractBusyThread.java:215) This tries to fix the problem; the problem should be monitored	2015-02-10 08:43:45 +01:00
Michael Peter Christen	1f5b5c0111	npe fix for latest scraper feature	2015-02-10 08:33:30 +01:00
Michael Peter Christen	ee97302a23	hack to make date detection faster (while it becomes a bit incomplete regarding language alternatives)	2015-02-09 18:46:06 +01:00
Michael Peter Christen	6578ff3ddb	enhanced suggest function	2015-02-09 18:45:07 +01:00
reger	fe6f5a395d	fix Umlaut handling in blekko heuristic search term http://mantis.tokeek.de/view.php?id=169 observation: blekko seams to block xxxbot agents (=0 results)	2015-02-08 23:40:33 +01:00
reger	23924348e2	url with semicolon or comma handling in proxy request apply patch supplied with bugreport http://mantis.tokeek.de/view.php?id=540	2015-02-07 22:01:54 +01:00
reger	9025fe3518	upd error message for proxy fix http://mantis.tokeek.de/view.php?id=539	2015-02-07 00:37:43 +01:00
Michael Peter Christen	97ba5ddbb7	configuration option for maxload limit for remote search	2015-02-04 01:12:25 +01:00
reger	c454ef69c6	add shortMemory check to heuristic search and skip operation on shortMemory (no request to remote openserch systems)	2015-02-03 03:08:34 +01:00
reger	9e1ec5fec4	refactor: just some more useages of constant for term ":[* TO *]"	2015-02-01 04:26:33 +01:00
reger	8c491f51a5	remove hardcoded initialization of language nav if not used	2015-02-01 00:29:28 +01:00
Michael Peter Christen	b5ac29c9a5	added a html field scraper which reads text from html entities of a given css class and extends a given vocabulary with a term consisting with the text content of the html class tag. Additionally, the term is included into the semantic facet of the document. This allows the creation of faceted search to documents without the pre-creation of vocabularies; instead, the vocabulary is created on-the-fly, possibly for use in other crawls. If any of the term scraping for a specific vocabulary is successful on a document, this vocabulary is excluded for auto-annotation on the page. To use this feature, do the following: - create a vocabulary on /Vocabulary_p.html (if not existent) - in /CrawlStartExpert.html you will now see the vocabularies as column in a table. The second column provides text fields where you can name the class of html entities where the literal of the corresponding vocabulary shall be scraped out - when doing a search, you will see the content of the scraped fields in a navigation facet for the given vocabulary	2015-01-30 13:20:56 +01:00
Michael Peter Christen	1cb290170e	refactoring of autotagging code (combined same code pieces)	2015-01-29 11:39:47 +01:00
Michael Peter Christen	c3b55455fc	enhanced initialization speed of vocabularies by using better normalization and by removal of unused data structures	2015-01-29 02:45:32 +01:00
Michael Peter Christen	68c605d637	replace with CommonPattern.SPACE for split	2015-01-29 02:28:03 +01:00
Michael Peter Christen	de3e373913	using precompiled CommonPattern.TAB for split	2015-01-29 02:22:28 +01:00
Michael Peter Christen	1f5047b15f	using precompiled pattern CommonPattern.SEMICOLON for splits	2015-01-29 02:19:41 +01:00
Michael Peter Christen	a8a2b7a803	persistency for vocabulary facet switch	2015-01-29 02:16:42 +01:00
Michael Peter Christen	efbc9a3561	introducting a new getConfig method which parses comma-separated llists from setting fields; refactoring for all places where such lists are parsed	2015-01-29 01:53:36 +01:00
Michael Peter Christen	69eacdf4eb	applying precompiled CommonPattern.COMMA.split to all places where split(",") was used	2015-01-29 01:46:22 +01:00
Michael Peter Christen	ac19690d30	refactoring with CommonPattern.COMMA	2015-01-29 01:35:28 +01:00
Michael Peter Christen	cf9b22ca5c	do not reindex based on vocabulary fields (there are meanwhile many of them) and some default settings	2015-01-29 01:22:28 +01:00
Michael Peter Christen	5a060c9f26	refactoring of reindexSolr (just replaced constant string)	2015-01-29 00:33:07 +01:00
Michael Peter Christen	b5a55c8b3d	fix for wkhtmltopdf (custom header does not work)	2015-01-28 17:45:25 +01:00
Michael Peter Christen	3d717b749a	fix for urlmaskfilter	2015-01-28 13:40:41 +01:00
Michael Peter Christen	bee5ee7cce	removed some warnings	2015-01-27 17:00:20 +01:00
Michael Peter Christen	783cf6fbc7	the LinkedBlockingQueue is much faster than the ArrayBlockingQueue (strange but this is the result of a test: ArrayBlockingQueue: 39461 lines / second; LinkedBlockingQueue: 60774 lines / second)	2015-01-27 16:53:09 +01:00
Michael Peter Christen	6390454652	fix for vocabulary on/off setting	2015-01-27 16:24:27 +01:00
Michael Peter Christen	a3c5995bde	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-01-26 14:13:17 +01:00
reger	5ca0762179	fix: eom on parsing ico file by genericImageParser trace: java.lang.OutOfMemoryError: Java heap space at java.awt.image.DataBufferInt.<init>(DataBufferInt.java:75) at java.awt.image.Raster.createPackedRaster(Raster.java:467) at java.awt.image.DirectColorModel.createCompatibleWritableRaster(DirectColorModel.java:1032) at java.awt.image.BufferedImage.<init>(BufferedImage.java:331) at net.yacy.document.parser.images.bmpParser$IMAGEMAP.<init>(bmpParser.java:149) at net.yacy.document.parser.images.bmpParser.parse(bmpParser.java:69) at net.yacy.document.parser.images.genericImageParser.parse(genericImageParser.java:116)	2015-01-24 23:17:07 +01:00
Michael Peter Christen	4cd2d68e03	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-01-24 07:10:47 +01:00
Michael Peter Christen	dc5700148f	update to latest code changes from json.org	2015-01-24 07:10:14 +01:00
reger	42b0672be3	Let auto-disabled crawls recover if low resource condition vanished. Analog to autodisabled DHT switch autodisabled crawls back on upon mem ok by remembering the autodisable by conf parameter.	2015-01-24 01:53:58 +01:00
Michael Peter Christen	287c528f46	replaced old JavaApplicationStub for Mac Application framework with new script. Adopted the YaCyApp environment and fixed a problem in the startYACY.sh application wrapper which caused wrong usage of logging option -l which caused that files had been written to the YaCy application folder. As a result of this fix, it is not necessary any more to change path settings in Info.plist if libraries are changed.	2015-01-23 11:30:13 +01:00
Michael Peter Christen	4c9d2a7c64	reverted 'do not show all options' strategy. This is actually confusing new users. Will be activated maybe again if there is an optional tutorial mode which can be switched on for this special purpose of running a tutorial.	2015-01-20 18:18:12 +01:00
Michael Peter Christen	7db2888336	fixed font size and print page generation in pdf snapshots	2015-01-20 17:14:14 +01:00
reger	24f68a4eb7	refactor opensearch heuristic introduce FederateSearchManager handling search heuristic to external systems via specific FederateSearchConnectors, which provide the query() functionallity, the translation to YaCy schema .toYaCySchema() and the search() routine to deliver results to searchevents, which is generally implemented in Abstract connector. The manager enforces now a min 15s delay between calls to external systems. Besides the OpensearchConnector a SolrFederateSearchConnector is available. It uses a additional config file for fieldname translation. default heuristicopensearch.conf: - openbdb.com removed - seems not longer to deliver results - config via solrconnector to datacite.org added (large technical library archive)	2015-01-19 03:30:35 +01:00
Michael Peter Christen	3b51636ecb	fix for mediawiki import	2015-01-12 00:35:47 +01:00
Michael Peter Christen	b07afbc115	a test with http://validator.w3.org/feed/#validate_by_input shows that the time format was wrong; we must use RFC-822	2015-01-09 16:45:43 +01:00
Michael Peter Christen	8cafdb989a	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-01-09 11:00:02 +01:00
reger	66839f73fa	remove debug limit from commit before	2015-01-09 02:52:18 +01:00
reger	4214f250d0	Add option for extended search (Autosearch) to Bookmark.html asking all connected peers for the searchterm added as description to the bookmark created by the bookmark icon. Intended for searches/research projects with not sufficient results from local and DHT selected remote target peers. Function: the process checks newly created bookmarks for description starting with "query=..." and takes this to ask every peer for 20 search results and adds it to the local index in a background job. link to start/stop the process added to /Bookmarks.html	2015-01-09 02:06:30 +01:00
reger	8e751d754a	- add javadoc to busythread with hint about the init parameter useage - remove obsolete 10_httpd config parameter	2015-01-09 01:31:57 +01:00
Michael Peter Christen	3e6c3e2237	documents pushed over the api/push_p.html interface will have their unique flag set by default	2015-01-06 15:22:59 +01:00
Michael Peter Christen	35c24608cc	fix for division by zero (rare cases)	2015-01-06 14:21:20 +01:00
Michael Peter Christen	4144c7cc52	do not write frame links to webgraph	2015-01-06 14:14:25 +01:00
reger	4eb89d7f15	revert clickservlet (default was indeed a mistakenly)	2015-01-05 09:10:20 +01:00
Michael Peter Christen	c9e2128260	please commit new files under your own name, this file was not created by me.	2015-01-05 08:18:19 +01:00
reger	d44d8996d0	Added a “don't store remote search results” option This is intended for peers who want to participate in the P2P network but don't wish to load/fill-up their index with metadata of every received search result. The DHT transfer is not effected by this option (and will work as usual, so that a peer disabling the new store to index switch still receives and holds the metadata according to DHT rules). Downside for the local peer is that search speed will not improve if search terms are only avail. remote or by quick hits in local index. To be able to improve the local index a Click-Servlet option was added additionally. If switched on, all search result links point to this servlet, which forwards the users browser (by html header) to the desired page and feeds the page to the fulltext-index. The servlet accepts a parameter defining the action to perform (see defaults/web.xml, index, crawl, crawllinks) The option check-boxes are placed in ConfigPortal.html	2015-01-04 11:10:45 +01:00
reger	c156548efe	add info text to metadata page (htmlresponsewriter) on no documents found	2015-01-04 02:59:21 +01:00
reger	3ac1d14a21	improve TexParser.mimeOf( fileextension ) by returning 1st defined in supported list. This prevents unusual mapping of supported fileextension -> mimetype (like htm=application/x-tex)	2015-01-02 04:20:02 +01:00
Michael Peter Christen	d2792a43fd	do not write iframe and embed links into webgraph, but use them anyway for crawling	2015-01-02 02:44:03 +01:00
Michael Peter Christen	3cd7deb3b8	do not flush non-errors to stdout because this is a concurrency issue. the flush-call appeared very often in thread dumps with high load, so this hopefully gives some performances	2014-12-28 15:48:37 +01:00
Michael Peter Christen	4e3e2acc69	Merge branch 'master' of gitorious.org:yacy/rc1-fixed_percent-encoding	2014-12-28 15:01:40 +01:00

... 5 6 7 8 9 ...

3765 Commits