yacy_search_server

mirror of https://github.com/yacy/yacy_search_server.git synced 2024-09-19 00:01:41 +02:00

Author	SHA1	Message	Date
Burkhard	4fdc11cae8	Update SearchEvent.java Fix NPE on disabled local SolrIndex, occuring on search moving to the 2nd result page. The debug purpose only setting to disabeling local SolrIndex (System Admin -> Debug Settings) should long term probably be removed from production code.	2017-02-22 02:01:48 +01:00
luccioman	cdc7f3e431	Switched some Solr fields from mandatory to optional These fields are default enabled but with no doubt not strictly mandatory with the current code base. As reported by @reger24, splitting between essential mandatory and optional fields is still to be improved to reflect the current YaCy needs.	2017-02-21 22:59:11 +01:00
luccioman	3475d8c1a9	Merge branch 'master' of https://github.com/yacy/yacy_search_server.git	2017-02-20 10:48:44 +01:00
luccioman	c68a8be2d9	Refactored and enforced Solr mandatory fields for proper operation - Added a new method to check activation of mandatory fields on Collection Configuration commit, consistently with checks previously performed in Switchboard startup and with mandatory fields in the default schema. - Reorganized default schema and CollectionConfiguration enumeration : moved no more mandatory fields in a specific section, and moved fields enabled at startup to the mandatory section. - Marked mandatory fields as required and with stronger font in the IndexSchema_p.html page	2017-02-20 10:48:07 +01:00
reger	334c70c37a	correct fromDate init value on missing param in api/timeline_p servlet revert test modification from last commit in AccessTracker.main	2017-02-20 00:14:14 +01:00
reger	cc770512d5	add hint of query syntax in AccessTracker log (qs=normal querystring, sq=solr-querystring) to allow to filter simple text queries for processing, remove toString for counter parameter use more predefined constants in solrservlet	2017-02-19 05:23:17 +01:00
luccioman	e5858bc8c8	Fixed a NullPointerException case possible on Index Export As reported by Palulukas in YaCy forum (http://forum.yacy-websuche.de/viewtopic.php?f=18&t=5944&sid=dcef5b899ab4aa9b40e3a3d158c13aed#p33454) the Index Export operation can fails, notably when the Solr index contains one or more documents with empty (despite required) "load_date_dt" field. This fixes the export failure when the situation finally occurs, but more should be done to harden verifications on minimum required fields.	2017-02-17 11:09:30 +01:00
reger	5e8879beb7	Reduce self generated content for text_t (visible text index field) to avoid repeat of tokenized url as description, continuation of `7e09bff4a1` `1409cabe8b` Add some javadoc, and not needed remove of omitted fields in postprocessing.	2017-02-16 01:43:14 +01:00
luccioman	1857651988	Added a new Debug/Analysis advanced settings subsection. As discussed in PR #93 with @JeremyRand and @reger24 this new advanced settings page includes: - a new setting to control remote Solr responses encoding - some existing debug settings which could not be set through the admin user interface	2017-02-09 11:05:06 +01:00
luccioman	526f2d6a8b	Fixed NPE case occurring when local solr index is disabled in search.	2017-02-09 10:59:41 +01:00
luccioman	08de58b6d3	Named a Thread without name for easier monitoring	2017-02-03 09:55:08 +01:00
reger	1f497ccad5	Add consistency check for related index fields upon load and save of index schema. To assemble the original link url for out-/inboundlinks, icons and pictures the _protocol_sxt and _urlstub_sxt is needed (due to the used data-reduced storage methode). Auto-enable _protocol_sxt if _urlstub_sxt is enabled. to be able to correctly assemble the original link url.	2017-01-28 00:36:03 +01:00
luccioman	68afe900d0	Added user-friendly controls over disk usage configuration settings. As mentioned in issue #103, control settings over YaCy disk usage already existed but lacked a user-friendly way to set them. I added it to the Performance_p.html administration page with a little refactoring on the "Resource Observer" fieldset for improved accessibility and HTML standards respect. Also added the possibility to enable/disable the autoregulation fonction from this page.	2017-01-27 15:47:15 +01:00
reger	95d2a28599	adjust the Field-Reindex Thread to verify and update the document id in case hash (ID) doesn't match document url (sku field).	2017-01-26 23:49:15 +01:00
luccioman	fc01b69eca	Fixed local image search pagination regression. As reported by @tglman on issue #90, when searching images on the local index only, pages next to the first were always empty. This was a regression from commit `c25e48e969`.	2017-01-25 09:54:39 +01:00
reger	581b00cc20	remove obsolete lastmodified calculation in WebgraphConfig	2017-01-17 23:45:56 +01:00
luccioman	0da1e6ba16	Factored code re-implementing DigestURL.hosthash() method. This ensure consistent implementation of the url host hash generation and easier usage finding in source code. Also added a unit test for this function.	2017-01-16 10:18:42 +01:00
luccioman	6a4d51d8f9	Cleaned up some Javadoc warnings.	2017-01-09 16:44:47 +01:00
luccioman	86dc198698	Fixed some JavaDocs broken links.	2017-01-09 09:57:53 +01:00
reger	4c9be29a55	fix concurrency issue with htmlParser using not current scraper data resulting in incorrect data for some html index metadata. Details see http://mantis.tokeek.de/view.php?id=717	2017-01-06 03:01:52 +01:00
reger	68d4dc5cc5	Complete harmonization RequestHeader getCookie with std ServletRequest to use javax.servlet.http.Cookie parameters. Depreciate now obsolete getHeaderCookies. Adjust setting of MaxAge to spec if >= 0 otherwise keep default.	2017-01-02 03:04:21 +01:00
reger	a1e5f7dbca	fix of fulltext.remove() by id of webgraph document webgraph has document hash in source_id_s	2017-01-01 23:53:44 +01:00
luccioman	1df558a6c6	Fixed YaCy proper shutdown triggered by SIGTERM signal. The main shutdown hook thread was not properly waiting for the main thread termination which consequently could not properly close resources and threads. After terminating a running YaCy peer this way (Ctrl+C in console, or kill <pid> for example), you could see the still existing DATA/yacy.running file. Tested with : - Debian Jessie openjdk 7 and 8 : regular shutdown, Ctrl+C, kill command, system restart while yacy is running - Windows 10 Oracle JDK 7 and 8 : non regression on regular shutdown	2016-12-28 09:47:27 +01:00
reger	b522d540b9	Include itemprop latitude/longitude (see schema.org) in attribute parsing for lat/lon. Harmonize number parsing for lat/lon to parseDouble. Fix endDate_dts value assignment.	2016-12-25 23:39:55 +01:00
luccioman	3ca695390c	FTP crawl start URLs : applied crawl profile depth control Applied rules : - when the FTP URL denotes a file resource, stack it as any start URL : eventually embedded links can be followed applying the usual depth rules - when the FTP URL denotes a directory, list files under this directory and stack them for crawl, and repeat the process on sub folders until crawl depth is reached	2016-12-22 16:25:09 +01:00
reger	8eb6fba59c	activate filetype navigator plugin and restrict config (append) of navs to not already actives. Dht results are now included in count this might over shoot on redundant dht and solr, while the previous solr facet based was always low.	2016-12-21 02:04:13 +01:00
luccioman	c25e48e969	Enabled displaying results after 14th page for local search queries. Fixes issue #90 for local queries only: Stealth mode, Portal mode or Intranet mode. For P2p mode, the issue would probably be difficult to solve with reasonable performance. This is still to dig. Also switched some InterreputedException catch log messages to warn level as this is normal behavior when shutting down a peer. Fixed yacysearch buttons navbar behavior to deal correctly with total results count or offset over 1000. Also improved the buttons navbar to be able to navigate over 10th page for local queries.	2016-12-20 14:52:33 +01:00
reger	bab4804d11	add FileTypeNavigator plugin	2016-12-19 23:56:03 +01:00
luccioman	467650c042	Hardened system update checks. When a downloaded archive release is corrupted, empty, or can not be opened for any reason, the update script must not be launched because it erases the existing lib/*.jar libraries.	2016-12-16 11:03:09 +01:00
luccioman	b5711b8fe1	Added some Javadocs.	2016-12-16 10:43:00 +01:00
reger	0758c868c9	add HostNavigator plugin	2016-12-13 22:14:16 +01:00
reger	60160877f5	bundle initialization of search navigation plugins in separate handler class to allow to use navigator map in config servlets (without need to create a search event)	2016-12-11 21:46:29 +01:00
luccioman	d27adc2b92	Fixed language detector initialization and NullPointerException cases. NullPointerException occurred when using and Identificator instance which encountered and error in its constructor. This error could be caused by a missing "langdetect" folder in the current folder of the main process, or by simultaneous first calls to the constructor, initializing concurrently the DetectorFactory.langlist. Fixes the mantis 714 (http://mantis.tokeek.de/view.php?id=714)	2016-12-05 18:12:21 +01:00
reger	f7e9f9be5f	move Digest auth checks from DefaultServlet to adminAuthenticated, eliminating the need to modify http header on Servlet container handled Digest authentication, to simulate Basic auth for YaCy servlets.	2016-11-29 03:20:33 +01:00
reger	44a6a4e795	fix authentication by hit in userdb (wrong parameter)	2016-11-24 00:16:22 +01:00
luccioman	aa9ddf3c23	Added control over Robots.txt active threads maximum number. When starting a crawl from a file containing thousands of links, configuration setting "crawler.MaxActiveThreads" is effective to prevent saturating the system with too many outgoing HTTP connections threads launched by the crawler. But robots.txt was not affected by this setting and was indefinitely increasing the number of concurrently loading threads until most ot the connections timed out. To improve performance control, added a pool of threads for Robots.txt, consistently used in its ensureExist() and massCrawlCheck() methods. The Robots.txt threads pool max size can now be configured in the /PerformanceQueus_p.html page, or with the new "robots.txt.MaxActiveThreads" setting, initialized with the same default value as the crawler.	2016-11-23 18:13:05 +01:00
reger	59130777a6	add high scored items first to YearNavigator (to make sure to be included in sorted view)	2016-11-22 01:17:33 +01:00
reger	08a0acc35d	make a YearNavigator availabel, useable as SearchEvent.naviator plugin. It can take any Date field of the index and displays a list of year strings in reverse order by the year (not the score/count). To allow to define the index field to use, the fieldname (and title can be appended to the navi's name "year" e.g. year:load_date_dt:LoadDate It works also with dates_in_content_dts field (from the graphical date navigator). Here the query parameter from: to: are used on selection as Query modifier (for other dates currently no query parameter available, so selection won't work to filter search results). Not included in the UI Searchpage layout config so far (for experiment with it manual change to conf needed).	2016-11-21 16:52:53 +01:00
reger	7742579ca4	make a LanguageNavigator availabel, useable for the SearchEvent.naviator plugin (not activated yet).	2016-11-21 00:19:11 +01:00
reger	bad8f87998	remove old/obsolete clear text "adminAccount" credential entry from init and setConfig (.,empty) from servlets/code	2016-11-20 00:20:47 +01:00
reger	811cf637f8	fix Jetty9YaCySecurityHandler, length check of Basic credential, add comment to SwitchboardConstants.AdminAccount const	2016-11-20 00:15:21 +01:00
reger	59448461d3	make use of userInRole for quick login verification	2016-11-16 01:42:53 +01:00
reger	2a4d826d9e	adjust servlet RequestHeader.getLocale init jvm defaultLocale matching UI language	2016-11-15 02:37:03 +01:00
reger	9db68acb4f	remove obsolete X_YACY... header declarations not in use (no writes, only remove and try to read). Obsolete parameter setupHttpClient	2016-11-14 22:53:26 +01:00
luccioman	84b81c1af0	Switched more URLs to relative ones when possible. This permits an easier and more flexible reverse proxy configuration. Some related mantis issues : http://mantis.tokeek.de/view.php?id=106 and http://mantis.tokeek.de/view.php?id=701	2016-11-08 03:05:51 +01:00
reger	8fe28a83f2	harmonize used lastmodified date for rwi and fulltext in storeDocument	2016-11-02 03:43:39 +01:00
reger	3d1d297308	refactor namespace navigator as part of navigatorplugin map, this allows the navigator to include counts all matches (rwi+fulltext). Fixing also unresolved_pattern in navigators title (of the counter) The use of inurl: query modifier as filter has not been changed keeping it as soft (unsharp) filter facet. Upd StringNavigator to prevent empty string form multivalued solr fields, removed date value conversion (better handled elsewhere, not need here).	2016-11-01 04:38:47 +01:00
reger	67f660523b	Make navigators underlaying indexfield name accessible in interface use interface in declaration and extend facet check to include navigator field.	2016-10-31 18:42:23 +01:00
reger	5eb3ee4e20	Add search navigator interface to allow for additional navigators (plugins) Prepared the first basic navigators (for authors and collections) for the list of SearchEvent.navigatorPlugins and adjusted servlet to use these. - this allows to configure display order of these navigators (by ordering config string) - eventually allows for additional and/or custom navigators using any available index field without need for changing servlets - the Collection navigation has been adjusted to exclude the internal, default robot_* and dht collections from displaying - rwi results are now also checked for navigatior by the refactored navi's So far no config options were added to customize or add navigators (may come later if route of upcoming modularization/plugin system is defined).	2016-10-31 02:17:43 +01:00
reger	fd3f58fcaa	improve query modifier parsing of "collection:" and possible collision with "on:" in case multiple collection modifier were entered (by mistake) http://mantis.tokeek.de/view.php?id=702	2016-10-31 00:43:01 +01:00
reger	af39a76bf6	Reduce number of default max. search navigator lines (from 10000) to 100 + make it configurable	2016-10-29 04:19:46 +02:00
reger	3c7220bc7b	Refacture rwi reference word position and word distance calculation used for rwi ranking. Main changes: - introduce a posintext() to access the stored value. This reduces also mem alloc of position array for WordReferenceRow (index access) - use the positions() array for joined references on multi-word queries if needed (otherwise allow positions() to be null - adjust assignments and the min() max() and distance() calculation accordingly	2016-10-23 19:40:02 +02:00
luccioman	f0639d810c	Customized name for Threads still using the default "Thread-n" pattern. This makes threads monitoring easier to read.	2016-10-22 17:17:21 +02:00
luccioman	7263d17436	Removed mentions of deprecated LURL-db. Thanks to LA_FORGE asking about if on YaCy forum ( http://forum.yacy-websuche.de/viewtopic.php?f=5&t=5895 )	2016-10-19 14:56:25 +02:00
reger	31d2a5645e	remove obsolete query variable leftover from `8fb370d9f8 (diff-1d4259005ebfddc11083387857a86175)` harmonize ranking shift parameter to 0xFF correct addresult weight parameter to long	2016-10-15 19:29:19 +02:00
luccioman	6e1959f469	Merge branch 'master' of https://github.com/yacy/yacy_search_server.git Conflicts: htroot/yacysearchitem.java source/net/yacy/cora/federate/solr/responsewriter/YJsonResponseWriter.java source/net/yacy/search/schema/CollectionConfiguration.java source/net/yacy/server/serverObjects.java	2016-10-14 11:29:55 +02:00
reger	685d8e86bf	Avoid frequent data type casting (float/long) for rwi score refactor to using long in URIMetadataNode too (and related call parameters) As remote rwi score's are not used (since v1.83) skip reading float-score , but keep in toString() for communication with older versions.	2016-10-14 01:17:34 +02:00
reger	e68b00678e	prevent negative score on URIMetadataNode - in the special case were no solr score is supplied. + assert before use & test case	2016-10-11 19:54:50 +02:00
luccioman	8d57b5b970	Added some javadocs.	2016-09-30 17:12:55 +02:00
luccioman	60df09fff9	Fixed some HTML validation errors : Illegal character in query Now encode space characters in URLs query part.	2016-09-30 10:54:53 +02:00
luccioman	b3b75b0498	Accessibility : add a customizable alternative text to YaCy log Applied W3C recommendations : https://www.w3.org/TR/html51/semantics-embedded-content.html#a-link-or-button-containing-nothing-but-an-image and https://www.w3.org/TR/html51/semantics-embedded-content.html#logos-insignia-flags-or-emblems	2016-09-22 16:08:33 +02:00
luccioman	3ee4f56c39	Improved ErrorCache behavior when switching networks Even after network switch, ErroCache was still holding a reference to the previous Solr cores, thus becoming useless until next YaCy restart. Initial error cache filling with recent errors from the index was also missing after the swtich.	2016-09-22 09:07:07 +02:00
luccioman	7d5ba2afa4	Added some JavaDoc and moved crawlStacker close at the right place.	2016-09-22 08:21:14 +02:00
luccioman	8edbcd8ad4	Log eventual Solr instances close errors. We do not want to block on this kind of error, but this should not silently fail as it may have later consequences.	2016-09-22 08:20:01 +02:00
reger	330768c8a2	fix for solr write.lock after mode change http://mantis.tokeek.de/view.php?id=686 The embedded core holds a lock on the index and must be closed. Earlier commit comment states that core should be closed with solr instance instead on close of connector. Adjusted the InstanceMirror.close() to take care of closing the embedded instance to release the lock. In 2 routines of fulltext this was already explicite implemented (disconnectLocalSolr). Now this disconnect is part of the InstanceMirror.close().	2016-09-22 00:16:22 +02:00
Michael Peter Christen	df51e4ef07	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2016-09-19 11:01:58 +02:00
Michael Peter Christen	e063aaf97f	enable fuzzy search, solr style (append a ~ to get a fuzzyness on the word)	2016-09-19 11:01:39 +02:00
reger	7f63fc50f3	prepare a IndexSegment test case for RWI index testing + prevent NPE in Segment.clear() on missing embedded solr instance.	2016-09-11 23:25:44 +02:00
luccioman	06d4f93d03	Merged master into postprocessing branch	2016-09-07 09:28:37 +02:00
reger	e310ec5f70	fix posInText ranking calculation to score 0 on no position info + fix Word posInText calc in Tokenizer to start with 1 + test case	2016-09-06 00:05:59 +02:00
reger	51c077f493	adjust the getTopics() and getTopicNavigator() to current useage - move the maxcount limit restriction completely to getTopicNavigator (as there not used in getTopics) - let search servlet use getTopics by default (w/o RWI connected check, as of now, Topics are available w/o any additional index interaction)	2016-09-05 00:07:01 +02:00
reger	cc2d9dd3f1	reactivate the use of included-in-topwords boost in postRanking + changed the postRanking to add one score only if word appears more as one time. + getTopics() unused code block rem'd (save performace)-> routine needs rework !	2016-09-04 00:09:45 +02:00
reger	6801673a07	apply postranking media search boost only on media queries	2016-09-03 03:37:40 +02:00
luccioman	8c49a755da	Postprocessing refactoring Added Javadocs to refactored methods. Added log warnings instead of silently failing some errors. Only fill collection1hosts when required ( shallComputeCR true).	2016-09-01 15:40:28 +02:00
luccioman	42f45760ed	Refactored postprocessing For easier understanding and performances profiling.	2016-08-31 12:16:25 +02:00
Michael Peter Christen	079112358c	Merge branch 'master' of https://github.com/yacy/yacy_search_server.git	2016-08-19 15:31:09 +02:00
Michael Peter Christen	efeb592661	don't do solr optimization, this create high IO load. We should leave this task to solr to do that on it's own instead of forcing it.	2016-08-19 15:30:53 +02:00
reger	4c7a77662a	eleminate dependency on file-extension in storeDocument but use supported mime-type to also support handling of urls w/o corresponding file-extension. For this refactor use of document.getParserObject() to alway return a Parser (for clean logic) and define/move the scraperObject as local var of AbstractParser. Adjust related calls to getParserObject (where actually a scraperObject is wanted). Addionally skip appending url token to parsed text for dht metadata entries (by default returned as result by rwi index).	2016-08-14 03:53:16 +02:00
reger	2910fe35c1	add missing scheduler calc of next exec_date (call of calculateAPIScheduler) - after last_exec_date is altered, next_exec_date should be recalculated - makes the recalculation of next_exec in advance (without api call surely made) in Switchbard.schedulerJob() obsolete Slightly modify next_exec calc. on missed event to now+schedule_time (from fix 10min)	2016-08-09 03:03:04 +02:00
reger	70d47ae38a	keep scheduler selection by repeat entry from `07311020d4` to allow exec schedule on actual exec event. Iterate on exec date (of advantage after interruption/shutdown) to schedule older or missed events first.	2016-08-08 02:19:48 +02:00
reger	7c3f932e5d	revert due to conflict with double count recording by schedulter / servlet by the commit under normal operation (no shutdown)	2016-08-08 01:57:31 +02:00
reger	07311020d4	postpone apicall exec date init until actual call fix for http://mantis.tokeek.de/view.php?id=677 The difference is on scheduling a large number of rss feeds and loading is not finished before shutdown of YaCy. The change makes sure not already loaded RSS will be loaded by the scheduler on next startup.	2016-08-07 05:08:55 +02:00
reger	fcad2d0744	add uses of config constant INDEX_RECEIVE_ALLOW	2016-07-27 02:16:20 +02:00
reger	35a7d57260	update lucenematchversion to current (5.2.0 -> 5.5.0) there should be no need for reindex by the update	2016-07-23 18:36:43 +02:00
luccioman	893a40995a	Merge branch 'master' of https://github.com/yacy/yacy_search_server.git	2016-07-04 21:24:40 +02:00
Michael Peter Christen	7466d390b2	small refactoring + do not accept too old peers during bootstrap	2016-07-04 11:02:15 +02:00
luccioman	6e96c7341a	Merge remote-tracking branch 'origin/master' Conflicts: htroot/Load_MediawikiWiki.java htroot/Load_PHPBB3.java htroot/ViewImage.java	2016-07-03 18:59:00 +02:00
reger	8d58a48029	remove wrong log line in CrawlSwitchboard + don't allow CrawlSwitchboard to exit application making network param unused	2016-07-02 20:33:23 +02:00
reger	b119ff65be	clean out not used Switchboard variables counter indexedPages, const xstackCrawlSlots	2016-06-14 01:50:32 +02:00
reger	bd8f7c11f5	Use transparent addToCrawler in AutoSearch instead of addToIndex This would likely also be of advantage for RSS import/schedule as following bug-reports suggest http://mantis.tokeek.de/view.php?id=569 http://mantis.tokeek.de/view.php?id=655	2016-06-01 01:14:22 +02:00
JeremyRand	433217b33e	Properly support multiple Boost Queries. (Previous code was broken because it concatenated multiple Boost Queries together rather than passing Solr an array.)	2016-05-20 20:17:51 -05:00
reger	d0a571bed2	del cytag trail for own index.html (save resource not used by default)	2016-05-19 01:59:00 +02:00
reger	7097dcbdbd	cleanup hack for partial Solr update on multivalued datefields has been fixed in Solr http://issues.apache.org/jira/browse/SOLR-8050	2016-05-06 02:47:04 +02:00
reger	f10ea3c155	clean-out unused SwitchboardConstants	2016-05-05 00:55:22 +02:00
reger	ef24593347	delete obsolete SEARCHRESULT busythread constants not used since 29.05.2013 18:27:27 `0c1a018bbd`	2016-05-04 01:30:10 +02:00
reger	6ecc180299	fix rwi doubledom return best (highest) ranking	2016-04-12 03:55:43 +02:00
reger	d9adc2c255	load handler for Transparent Proxy on startup only if feature is activated to save the resources and keep handler chain small if the feature is not used. +add a warning message on settingsack_p page to restart on first activation	2016-03-25 05:26:48 +01:00
Michael Peter Christen	b89465d952	0N - basic dump upload servlet infrastructure, to share index dumps within an experimental new sharing model	2016-03-11 18:12:13 +01:00
Michael Peter Christen	849ab671a9	0n: modified the p2p bootstraping process - rules had been too tight and did not support the re-start of a network with just one principal peer.	2016-03-11 08:54:42 +01:00
Michael Peter Christen	a6bf0b1649	0N - added option to generate index export files for a specific number of minutes in the past and reverted latest change. The export file dump will now contain four data elements: f - first date of index entry write date, l - last date of index write date, n - now-date of index dump time, c - count of numbers inside the dump. '0N' denotes a series of changes which will lead to the opportunity to exchange index data dumps in a way that is needed to integrate ZeroNet index data. This will be based on index dump sharing; that causes this commit.	2016-02-23 18:56:20 +01:00
reger	06d0e2aeb9	result heuristic (also used in greedy learning mode) to use outbound links if result is full index doc. Otherwise use default loader methode. - Above brought up that parser start url parameter, declared as AnchorURL uses only methodes of parent object DigestURL (changed parameter declaration accordingly).	2016-02-16 02:05:58 +01:00
reger	caf9e98f09	put metadata dc_publisher in corresponding schema field	2016-02-14 21:13:25 +01:00
luc	3f338777f7	Also check and index eventual icon url information from metadata.	2016-02-11 09:33:20 +01:00
reger	6f0b073bf3	override detected language (statistic langdetect) only with TLD determided language if langdetect probability is not high. + additionally truncate zh-cn / zh-tw returned by langdetect to 2 char ISO639-1 zh used by YaCy	2016-02-07 21:16:22 +01:00
luc	07222b3e1a	Added favicon url transmission in RWI chunks.	2016-02-05 17:05:36 +01:00
luc	480772c070	Fixed json search results from commit "Improved URLLicence reliability"	2016-02-05 15:23:29 +01:00
reger	535d4bf75f	respect hidden attribute for file and smb directory listing (hidden directories are not listed, effects crawling of local file system)	2016-02-04 19:16:00 +01:00
luc	3cc5619d93	Improved HTML icons indexing and rendering in search results. See http://mantis.tokeek.de/view.php?id=629	2016-02-02 09:57:54 +01:00
reger	a6617ad887	expand initRemoteCrawler() to terminate worker threads if called to deactivate remote crawl. On startup we save the resources for remote crawler if disabled. Once started threads are running idle after disable remote crawl. Now threads are terminated to save the resources also while disabeling during runtime. + remove empty class Channels	2016-01-28 23:14:09 +01:00
reger	ed3e16e092	apply remote result count config value to Bookmark Autosearch + prepare to make the widely unused Bookmark feature optional	2016-01-15 02:10:10 +01:00
Ryszard Goń	a98c395023	Add the Autocrawl thread	2016-01-14 00:50:23 +01:00
Ryszard Goń	1728cd30c6	Create autocrawl profiles	2016-01-12 16:28:34 +01:00
luc	571bc55937	Refactoring : use StandardCharsets constants instead of hard-coded charset names.	2016-01-05 23:37:05 +01:00
reger	1af0e9ef74	remove workaround for Solr bug regarding multivalued date fields fixed in 5.4.0 http://issues.apache.org/jira/browse/SOLR-8050	2016-01-03 01:11:27 +01:00
reger	a58d34a4e8	check error URL cache before adding errorDoc to index - del obsolete related switchboardconstant	2016-01-02 05:03:57 +01:00
reger	cd26717ba2	fix low memory status hint (dht-in disabled) http://mantis.tokeek.de/view.php?id=619	2015-12-29 20:38:45 +01:00
sixcooler	dce1cb65c4	Merge remote-tracking branch 'choose_remote_name/master'	2015-12-28 23:20:42 +01:00
reger	6d54eb3d36	skip loading document on crawl start for YMark bookmarks by adding a constructor giving the already loaded document as parameter.	2015-12-26 01:15:07 +01:00
reger	45b9bd8403	adjust MultiProtocolURL.protocol detection to handle mailto with "://" in parameters, and feeding hyperlinks to webgraph processing.	2015-12-21 04:42:26 +01:00
reger	dec3e6ad96	fix: adjust urlstub for mailto links (skip protocol)	2015-12-19 20:13:33 +01:00
luc	8c4ab9c76b	Added an option to eventually limit size of remote solr documents put to local index. See mantis #626.	2015-12-16 02:20:03 +01:00
reger	28b8bc290a	fix use of NETWORK_SEARCHVERIFY for rwi verification was not used to set the searchevent parameter (done in SearchEventCache.getEvent) - remove unused corresponding QueryParams.filterfailurls param.	2015-12-13 20:01:49 +01:00
reger	020630efd8	remove unused network scanner parameter from queryparameter Search event is not using networkscanner (removed filterscannerfail param always init to false)	2015-12-13 02:50:08 +01:00
luc	ad5586f8f6	Merge branch 'master' of https://github.com/yacy/yacy_search_server	2015-12-08 03:35:36 +01:00
luc	8ebefa4233	Fixed MediaWiki import : DCEntry conversion to SolrInputDocument was failing. Looks like it was broken since Commit `b43811d38c`	2015-12-08 03:34:03 +01:00
reger	cdb8f3b10d	make current ranking score value avail. to search interface / api Update the result score result field with the result queue ranking value to reflect the actual calculated/used score, for rwi & solr stack results. (calc. etc. is unchanged, it's just that result entry carries the latest val as api retrieves the number from it)	2015-12-08 03:17:32 +01:00
Michael Peter Christen	ef8cd80593	fix for npe	2015-12-03 00:33:13 +01:00
reger	0347bfa71f	Apply collection query constraint/modifiert to rwi result stack. Collection is not available in pure rwi entries (but in local solr metadata) But if user wishes to filter by query constraint also rwi shall adhere to this (even if only rwi entries with parsed or solr received metadata may fit)	2015-12-02 22:57:59 +01:00
reger	ca3d26a401	harmonize wordsintitle & CollectionSchema.title_words_val calculation, remove obsolete partial init of wordreference from urimetadata	2015-11-15 06:06:37 +01:00
reger	52a9040ae6	Sort out double keywords (dc_subject) early in parsed documents - by direct using Set vs. List - remove not neede String[] getter	2015-11-13 01:48:28 +01:00
sixcooler	646afe9183	do not store subfield *_coordinate + make all num-fields being docvalues	2015-11-10 20:45:33 +01:00
sixcooler	194df613de	not using 'location' as defaultfacetfield - since we removed it being default.	2015-11-10 20:43:58 +01:00
sixcooler	4a905ec134	fix to not let the AccessTracker-Log grow to much, but have enough data to monitor. (+gitignore-correction)	2015-11-10 20:27:17 +01:00
reger	a60b1fb6c2	differentiate api call getLocalPort() from getConfigInt()	2015-10-31 23:09:03 +01:00
reger	11f3666660	increase use of pre.defined CATCHALL_QUERY string	2015-10-31 19:44:31 +01:00
reger	a58ee49307	Optimize internal imagequery focus on using content_type to select images (in favor of url file extension)	2015-10-31 19:18:46 +01:00
Michael Peter Christen	151ccd50a9	fix for image size field values (must be multi-valued)	2015-10-14 15:16:16 +02:00
reger	43c27aa550	upd to solr/lucene 5.3.1	2015-10-03 23:20:33 +02:00
Michael Peter Christen	3d7dd9d3aa	follow-up to latest commit: also flush the search cache if all crawls had been terminated.	2015-10-01 13:21:28 +02:00
Michael Peter Christen	c737ff235d	in case that the include_string contains several entries including 1-char tokens and also more-than-1-char tokens, then remove the 1-char tokens to prevent that we are to strict. This will make it possible to be a bit more fuzzy in the search where it is appropriate.	2015-10-01 13:09:33 +02:00
reger	7889fc2389	Hack to prevent Solr issue on partial update on a document containing multivalued date field (regardless if these fields part of update). Switch partial update option off in postprocessing if schema contains *_dts (multivalued date field). see http://mantis.tokeek.de/view.php?id=601	2015-09-13 20:23:15 +02:00
reger	3428b6f13b	improve filtering by filetype navigator. The used url-filter for filetype doesn't require ".ext" resulting in too many matches, add a sort-out filter for RWI results.	2015-09-07 02:36:22 +02:00
reger	e37a4f0b3d	prevent metadata records in index w/o valid url by throwing MalformedURL exception on URIMetadataNode creation	2015-09-06 22:19:05 +02:00
reger	802ccaead6	fix init of error cache, use latest faildates => load_date_dt	2015-09-02 02:36:31 +02:00
reger	dba7f15073	apply same size constrain on result image from doc as for linked images see `19f1308bf0`	2015-09-01 23:22:48 +02:00
sixcooler	87e4abe393	fight the fieldcache by usind DocValues: in Solr-5.x the fieldcache has moved and was not cleared anymore. This results in an huge fieldcache. (http://lucene.apache.org/#highlights-of-the-lucene-release-include https://issues.apache.org/jira/browse/LUCENE-5666) Here I try to use DovValues where it is possible. For this I used the Api-Scheme as new basis für the Solr-Schema. This needs at least a complete optimization of the Solr-Index to get a smaller FieldCache. Everything that is indexed with these setting will not use the Fieldcache at all.	2015-08-31 20:24:41 +02:00
reger	eaf0e8ff2c	start recording/indexing pixel size for image document as for linked images	2015-08-31 01:58:36 +02:00
reger	c33229fc0c	check mime prior to ext for metadata modification for images	2015-08-30 23:02:19 +02:00
reger	19f1308bf0	enforce th result images limit to > 16x16px for linked images http://mantis.tokeek.de/view.php?id=594	2015-08-30 02:19:52 +02:00
Michael Peter Christen	8028410ab7	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-08-10 14:27:53 +02:00
Michael Peter Christen	df3314ac1a	added a new facet type based on a probabilistic classifier using bayesian filters. This can be used to classify documents during indexing-time using a pre-definied bayesian filter. New wordings: - a context is a class where different categories are possible. The context name is equal to a facet name. - a category is a facet type within a facet navigation. Each context must have several categories, at least one custom name (things you want to discover) and one with the exact name "negative". To use this, you must do: - for each context, you must create a directory within DATA/CLASSIFICATION with the name of the context (the facet name) - within each context directory, you must create text files with one document each per line for every categroy. One of these categories MUST have the name 'negative.txt'. Then, each new document is classified to match within one of the given categories for each context.	2015-08-10 14:27:44 +02:00
reger	1409cabe8b	exclude more default search fields from text copy to text_t for metadata index documents	2015-08-09 21:01:30 +02:00
Michael Peter Christen	dbbad23e12	removed warnings	2015-08-03 05:37:34 +02:00
Michael Peter Christen	c14bc8d9b7	revert of fq transformation (recent fix)	2015-08-03 05:15:34 +02:00
Michael Peter Christen	11a848da5a	Merge branch 'master' of git@github.com:yacy/yacy_search_server.git	2015-08-02 14:53:36 +02:00
Michael Peter Christen	b94bd7f20a	a collection of search query enhancements: - fixed superfluous space in query field list - fixed filter query logic - removed look-ahead query which caused that each new search page submitted two solr queries - fixed random solr result orders in case that the solr score was equal: this was then re-ordered by YaCy using the document hash which came from the solr object and that appeared to be random. Now the hash of the url is used and the score is additionally modified by the url length to prevent that this particular case appears at all.	2015-08-02 14:52:41 +02:00
reger	cb67eb7baf	use more absolute path for config file opening as suggested in pull request 5 (https://github.com/yacy/yacy_search_server/pull/5)	2015-08-01 23:54:26 +02:00
Michael Peter Christen	de8cfbe1d7	added export option to export the fulltext of the search index text only	2015-07-30 03:21:40 +02:00
Michael Peter Christen	0aa6fcf259	remove old vocabularies and synonyms before adding new	2015-07-10 16:47:19 +02:00
reger	f91298d3b6	fix one implicit Integer/Long type conversion -> causes Java 1.8 compile error	2015-07-08 03:02:10 +02:00
reger	821262a179	add CommonPattern for multiple spaces to eliminate empty split words on following spaces	2015-07-04 22:49:01 +02:00
Michael Peter Christen	90f75c8c3d	added enrichment of synonyms and vocabularies for imported documents during surrogate reading: those attributes from the dump are removed during the import process and replaced by new detected attributes according to the setting of the YaCy peer. This may cause that all such attributes are removed if the importing peer has no synonyms and/or no vocabularies defined.	2015-07-02 00:23:50 +02:00
Michael Peter Christen	593de05922	enhanced surrogate import process speed (dramatically!)	2015-06-29 12:28:34 +02:00
Michael Peter Christen	694b22f165	migration to Solr 5.2: huge benefits - this is a lot faster! This is a very complex migration: many classes had been renamed or removed, dependencies changed and the solr index type is now aligned to be a solr cloud repository. Together with the Solr 5.2 library update, one other dependent library had been updated as well: httpclient 4.4->4.4.1 Older indexes are migrated from 4_10 to 5_2. However, the new index structure is more efficient and we recommend to re-index everything. Please use the index export before you do the update to a large surrogate xml file. After the update, start with an empty index and then initialize this with your dump.	2015-06-24 01:55:51 +02:00
reger	0fab445b19	Resourceobserver log warning - deleting releases files - only on actual deletes instead of entering routine	2015-06-10 02:35:37 +02:00
reger	c973f94936	add log entry on release file delete by ResourceObserver	2015-06-08 03:17:12 +02:00
reger	121972752c	implement deleteOldDownloads in RexourceObserver on low diskspace - direct assign sb.observer (skip redundant InitThread)	2015-06-08 02:52:13 +02:00
reger	49b79987c9	remove obsolete searchfl work table was used to register urls with not complete words in snippet but is never accessed	2015-06-04 22:44:01 +02:00
Michael Peter Christen	d0aff91f23	fix for index import	2015-06-01 01:56:09 +02:00
Michael Peter Christen	34de1e8cbc	gzip compression will perform more efficient and with better compression level	2015-06-01 01:24:33 +02:00
Michael Peter Christen	98be59ce9c	full solr xml exports will now be automatically compressed during export. That makes it possible to export a solr xml dump even if disc space is low.	2015-05-30 19:02:54 +02:00
Michael Peter Christen	b43811d38c	added surrogate import process for exported solr dumps. Just throw your solr dump file into DATA/SURROGATES/in/ and it will be imported!	2015-05-30 13:19:59 +02:00
Michael Peter Christen	c7576d6028	added a full solr export to the IndexControlURLs_p.html servlet. The export function is also now the default export option. The export file format for a full solr export is very similar to a solr search result xml, only the <lst name="responseHeader"> tag is missing. The exported xml has a special line termination feature: all documents will be exported into a single line without any CR in between. That means that every document is completely inside a single line. While this is not readable at all for humans, it is very useful for linux line processing scripts, like grep. Using grep it will be easy to select single documents which match for a given pattern. Such dumps shall be importable with the DATA/SURROGATE/in import function, but that import is not yet adopted to the new file format.	2015-05-29 15:05:52 +02:00
Michael Peter Christen	197f7449e5	All entities of crawl profiles are now editable in the crawl profile editor.	2015-05-28 16:07:40 +02:00
reger	1d8e1e4bac	- Image search expand box, adjust javascript hs padtominsize parameter, to make sure expand box doesn't shrink on small images - asure ImageResult.imagetext has value for the link text (use filename if no alt text given)	2015-05-27 02:31:13 +02:00
reger	af57fbefad	use available mime (instead null) on imageresult from metadatanode	2015-05-26 23:54:04 +02:00
reger	000dde9511	Eleminate duplication of values for search ResultEntry by instatiation from URIMetadataNode, by eleminating differentiation of ResultEntry/URIMetadataNode. - moved remaining ResultEntry functionallity to URIMetadataNode - for 1:1 functionallity added a function makeResultEntry() - removed ResultEntry - refactored related code Main difference is after makeResultEntry the text_t content is removed and alternative title/url strings for display are calculated. Main difference left is, that	2015-05-26 04:15:00 +02:00
reger	29c4aa3991	fix compiler notification of missing serialID from last commit	2015-05-25 21:51:32 +02:00
reger	3d53da8236	refactor ResultEntry to be based on MetadataNode/SolrDocument to share/reuse common access routines	2015-05-25 21:28:48 +02:00
reger	d882991bc5	Implement sharing of ioDispatcher for term & citation index as proposed in ioDispatcher description	2015-05-25 19:46:26 +02:00
reger	370ba9da71	On imageSearch prefere mime to sort out none-image documents Generalize the hack to prevent urls with just a img extension beeing returned improving http://mantis.tokeek.de/view.php?id=528	2015-05-24 21:48:58 +02:00
reger	3e742d1e34	Init remote crawler on demand If remote crawl option is not activated, skip init of remoteCrawlJob to save the resources of queue and ideling thread. Deploy of the remoteCrawlJob deferred on activation of the option.	2015-05-23 02:06:39 +02:00
reger	f3ce99bfb8	fix extract of inboundlinks_protocol_sxt url counter maybe > 999	2015-05-14 00:03:09 +02:00
reger	2bc9cb5828	fix early return in addToCrawler check / handle all supplied urls after error url	2015-05-13 21:58:43 +02:00
Michael Peter Christen	0710648c31	enable api calls with very long urls	2015-05-11 14:42:21 +02:00
reger	1481a8ab56	add opensearch rss results to dht collection (due to text = snippet) which is used to differentiate meta from full data - make sure check for dht is not dependant on number of collection entries	2015-05-10 18:52:33 +02:00
reger	752eec6697	fix NPE in addToIndex when used outside searchEvent	2015-05-10 05:18:23 +02:00
Michael Peter Christen	ff29b0e503	added option to re-index exported xml snapshot dumps to HTCACHE/snapshots by just placing them in the SURROGATES/in path	2015-05-08 15:30:26 +02:00
Michael Peter Christen	6f4fe4b175	revert of `8a7c68e4c7` keeping surrogates after processing is essential for some users. If the space they are taking is too high, please set up an automatic deletion process (like a cronjob).	2015-05-08 14:01:30 +02:00
Michael Peter Christen	97930a6aad	added must-not-match filter to snapshot generation. also: fixed some bugs	2015-05-08 13:46:27 +02:00
Michael Peter Christen	9d8f426890	adding a try-catch to link graph processing to prevent that a single malformed url interrupts the storage process	2015-05-08 10:38:33 +02:00
reger	8a5b8f8789	on bookmaring of search result, remember orig. query in separate bookmark property (instead of using the description field) - adjust display and autosearch - don't overwrite existing bookmark but combine info	2015-05-03 02:31:50 +02:00
reger	7224209486	break out of NormalizeDistributor loop on timeout	2015-05-02 02:36:18 +02:00
reger	47e61f8325	fix typo in image filter query (extra bracket)	2015-04-28 03:12:14 +02:00
reger	4b4ab6799f	fix String out of range in Collection Nav see http://mantis.tokeek.de/view.php?id=573	2015-04-27 22:38:40 +02:00
reger	5408448a56	skip redundant add. of keywords to text search uses keywords as default search field	2015-04-17 02:14:13 +02:00
reger	296e97c78e	put https port in peers dna as we flag if a peer is accesible via https, we need to know the port if we want to use is (e.g. for interYaCy communication) start to provide / tansport the port by recording it in peers dna. - add https link on the Network.html lock symbol	2015-04-16 02:36:12 +02:00
Michael Peter Christen	fed26f33a8	enhanced timezone managament for indexed data: to support the new time parser and search functions in YaCy a high precision detection of date and time on the day is necessary. That requires that the time zone of the document content and the time zone of the user, doing a search, is detected. The time zone of the search request is done automatically using the browsers time zone offset which is delivered to the search request automatically and invisible to the user. The time zone for the content of web pages cannot be detected automatically and must be an attribute of crawl starts. The advanced crawl start now provides an input field to set the time zone in minutes as an offset number. All parsers must get a time zone offset passed, so this required the change of the parser java api. A lot of other changes had been made which corrects the wrong handling of dates in YaCy which was to add a correction based on the time zone of the server. Now no correction is added and all dates in YaCy are UTC/GMT time zone, a normalized time zone for all peers.	2015-04-15 13:17:23 +02:00
Michael Peter Christen	b060ba900d	added parsing of contentprop attribute in html tags for content='startDate' and content='endDate'. The value of these field is now written to new solr fields startDates_dts and endDates_dts.	2015-04-13 16:20:00 +02:00
Michael Peter Christen	4cb4f67f38	added parsing of dd, dt and article html fields. The parsed result is written to special solr fields which are deactivated by default.	2015-04-12 22:02:45 +02:00
reger	1395f10e95	fix typecast for css links	2015-04-12 01:11:47 +02:00
Michael Peter Christen	abaaaef5f1	fix for filter queries	2015-04-11 12:30:29 +02:00
Michael Peter Christen	f5a032f293	split query into filter query and text query to get better ranking results and faster results	2015-04-07 16:10:13 +02:00
Michael Peter Christen	2e88028c1a	when selecting collections in navigation, do show the un-selected collections in search result. When selecting one of them in another search, switch off the previously selected collection. This actually turns the collection navigation modifier into a radio-button like behaviour	2015-04-07 13:13:58 +02:00
Michael Peter Christen	fa7edc9f7a	refactoring of filter queries (several queries instead only one)	2015-04-02 13:27:47 +02:00
Michael Peter Christen	40389987ec	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-04-01 18:18:05 +02:00
Michael Peter Christen	f9ba50379d	added an expansion option to search facets on result page: - if less or equal of 8 facet options are present, they are shown by default - if more facet options are present, they are hidden To view or hide all facets, just click on the facet header bar	2015-04-01 18:17:52 +02:00
reger	1f0f77bb77	make location facet return results for location nav facet of field coordinate_p does not return results, now using coordinate_p_0_coordinate as alternative to get facet counts. As the actual facet value is not used this should not harm any analysis (even if facet is a incomplete location). If facet value is used in future likely *_geohash field could be introduced (for facet and other ... as transport value)	2015-04-01 01:57:56 +02:00
Michael Peter Christen	9bf0d7ecb9	added a new collection type 'dht' to all documents from the peer-to-peer interface to distinguish rich and poor document data. This also reverts some changes from commit `796770e070` because the firstSeen database is the wrong method to distinguish these types of data	2015-03-24 12:32:39 +01:00
reger	f63fff9008	fix snippet containig number with comma as desmo point http://mantis.tokeek.de/view.php?id=344 to keep it as one word (by altering the split regex) - added sniipet test case with number - regex for word split to match multiple splitcars	2015-03-16 02:03:40 +01:00
reger	b241264632	fix error on *abc query input http://mantis.tokeek.de/view.php?id=486	2015-03-15 22:31:47 +01:00
reger	7e09bff4a1	exclude default search fields from text copy to text_t for metadata index documents (reduce text redundance)	2015-03-08 21:49:23 +01:00
reger	8af70950d9	harmonize snippet computation to considere description_txt always (solr hl & internal). For now just added desc to text list for computation, could be further equalized with hl computation.	2015-03-05 02:22:05 +01:00
Michael Peter Christen	fd4e2c809a	Show dates in the content of a document in the search result: - if an eventDate is given in the search result, replace the document date with the event date and prefix it with the string "on ". - the document date is omitted if a date from the cent is shown Added also the date as fields in the json and rss result sets.	2015-03-02 18:00:20 +01:00
Michael Peter Christen	d9d3111d10	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-03-02 04:31:05 +01:00
Michael Peter Christen	535f1ebe3b	added a new way of content browsing in search results: - date navigation The date is taken from the CONTENT of the documents / web pages, NOT from a date submitted in the context of metadata (i.e. http header or html head form). This makes it possible to search for documents in the future, i.e. when documents contain event descriptions for future events. The date is written to an index field which is now enabled by default. All documents are scanned for contained date mentions. To visualize the dates for a specific search results, a histogram showing the number of documents for each day is displayed. To render these histograms the morris.js library is used. Morris.js requires also raphael.js which is now also integrated in YaCy. The histogram is now also displayed in the index browser by default. To select a specific range from a search result, the following modifiers had been introduced: from:<date> to:<date> These modifiers can be used separately (i.e. only 'from' or only 'to') to describe an open interval or combined to have a closed interval. Both dates are inclusive. To select a specific single date only, use the 'to:' - modifier. The histogram shows blue and green lines; the green lines denot weekend days (saturday and sunday). Clicking on bars in the histogram has the following reaction: 1st click: add a from:<date> modifier for the date of the bar 2nd click: add a to:<date> modifier for the date of the bar 3rd click: remove from and date modifier and set a on:<date> for the bar When the on:<date> modifier is used, the histogram shows an unlimited time period. This makes it possible to click again (4th click) which is then interpreted as a 1st click again (sets a from modifier). The display feature is NOT switched on by default; to switch it on use the /ConfigSearchPage_p.html servlet.	2015-03-02 04:30:10 +01:00
reger	d7259419f3	postpone raw snippet html encoding upon use instead of during init of snippet adressing http://mantis.tokeek.de/view.php?id=551	2015-02-28 19:02:18 +01:00
reger	9b0de2de64	introduce getQueryFields to return default query fields (queryparamter QF) calculated from boostfields config, making sure title, description, keywords and content is always searched. - apply change to solrServlet makes sure every remote query uses at least all locally defined boost fields for search - apply to local solr search - simplify select query by using QF defaults	2015-02-23 23:12:07 +01:00
reger	4b97ddb9ec	stop sending crawl receipts if receiver got offline	2015-02-17 03:16:10 +01:00
reger	fba34e12ef	fix formatting issue if snippet contains html code replacement for reverted commit `61f42a7928`	2015-02-15 20:39:20 +01:00
reger	e48720a58c	fix NPE in snippet computation	2015-02-15 05:30:14 +01:00
Michael Peter Christen	97ba5ddbb7	configuration option for maxload limit for remote search	2015-02-04 01:12:25 +01:00
reger	9e1ec5fec4	refactor: just some more useages of constant for term ":[* TO *]"	2015-02-01 04:26:33 +01:00
reger	8c491f51a5	remove hardcoded initialization of language nav if not used	2015-02-01 00:29:28 +01:00
Michael Peter Christen	b5ac29c9a5	added a html field scraper which reads text from html entities of a given css class and extends a given vocabulary with a term consisting with the text content of the html class tag. Additionally, the term is included into the semantic facet of the document. This allows the creation of faceted search to documents without the pre-creation of vocabularies; instead, the vocabulary is created on-the-fly, possibly for use in other crawls. If any of the term scraping for a specific vocabulary is successful on a document, this vocabulary is excluded for auto-annotation on the page. To use this feature, do the following: - create a vocabulary on /Vocabulary_p.html (if not existent) - in /CrawlStartExpert.html you will now see the vocabularies as column in a table. The second column provides text fields where you can name the class of html entities where the literal of the corresponding vocabulary shall be scraped out - when doing a search, you will see the content of the scraped fields in a navigation facet for the given vocabulary	2015-01-30 13:20:56 +01:00
Michael Peter Christen	68c605d637	replace with CommonPattern.SPACE for split	2015-01-29 02:28:03 +01:00
Michael Peter Christen	a8a2b7a803	persistency for vocabulary facet switch	2015-01-29 02:16:42 +01:00
Michael Peter Christen	69eacdf4eb	applying precompiled CommonPattern.COMMA.split to all places where split(",") was used	2015-01-29 01:46:22 +01:00
Michael Peter Christen	5a060c9f26	refactoring of reindexSolr (just replaced constant string)	2015-01-29 00:33:07 +01:00
Michael Peter Christen	3d717b749a	fix for urlmaskfilter	2015-01-28 13:40:41 +01:00
Michael Peter Christen	bee5ee7cce	removed some warnings	2015-01-27 17:00:20 +01:00
reger	42b0672be3	Let auto-disabled crawls recover if low resource condition vanished. Analog to autodisabled DHT switch autodisabled crawls back on upon mem ok by remembering the autodisable by conf parameter.	2015-01-24 01:53:58 +01:00
Michael Peter Christen	7db2888336	fixed font size and print page generation in pdf snapshots	2015-01-20 17:14:14 +01:00
reger	24f68a4eb7	refactor opensearch heuristic introduce FederateSearchManager handling search heuristic to external systems via specific FederateSearchConnectors, which provide the query() functionallity, the translation to YaCy schema .toYaCySchema() and the search() routine to deliver results to searchevents, which is generally implemented in Abstract connector. The manager enforces now a min 15s delay between calls to external systems. Besides the OpensearchConnector a SolrFederateSearchConnector is available. It uses a additional config file for fieldname translation. default heuristicopensearch.conf: - openbdb.com removed - seems not longer to deliver results - config via solrconnector to datacite.org added (large technical library archive)	2015-01-19 03:30:35 +01:00
Michael Peter Christen	3b51636ecb	fix for mediawiki import	2015-01-12 00:35:47 +01:00
Michael Peter Christen	8cafdb989a	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2015-01-09 11:00:02 +01:00
reger	66839f73fa	remove debug limit from commit before	2015-01-09 02:52:18 +01:00
reger	4214f250d0	Add option for extended search (Autosearch) to Bookmark.html asking all connected peers for the searchterm added as description to the bookmark created by the bookmark icon. Intended for searches/research projects with not sufficient results from local and DHT selected remote target peers. Function: the process checks newly created bookmarks for description starting with "query=..." and takes this to ask every peer for 20 search results and adds it to the local index in a background job. link to start/stop the process added to /Bookmarks.html	2015-01-09 02:06:30 +01:00
Michael Peter Christen	3e6c3e2237	documents pushed over the api/push_p.html interface will have their unique flag set by default	2015-01-06 15:22:59 +01:00
reger	4eb89d7f15	revert clickservlet (default was indeed a mistakenly)	2015-01-05 09:10:20 +01:00
reger	d44d8996d0	Added a “don't store remote search results” option This is intended for peers who want to participate in the P2P network but don't wish to load/fill-up their index with metadata of every received search result. The DHT transfer is not effected by this option (and will work as usual, so that a peer disabling the new store to index switch still receives and holds the metadata according to DHT rules). Downside for the local peer is that search speed will not improve if search terms are only avail. remote or by quick hits in local index. To be able to improve the local index a Click-Servlet option was added additionally. If switched on, all search result links point to this servlet, which forwards the users browser (by html header) to the desired page and feeds the page to the fulltext-index. The servlet accepts a parameter defining the action to perform (see defaults/web.xml, index, crawl, crawllinks) The option check-boxes are placed in ConfigPortal.html	2015-01-04 11:10:45 +01:00
Michael Peter Christen	d2792a43fd	do not write iframe and embed links into webgraph, but use them anyway for crawling	2015-01-02 02:44:03 +01:00
Michael Peter Christen	ecb6a59e9e	do not translate gif images into png images for thumbnails. Instead, stream the original to the search result thumb viewer. This has two reasons: - animated gifs cause 100% cpu and deadlocks in the jvm gif parser; a known bug which is obviously not yet fixed - animated gifs now appear in the search result also as animation	2014-12-28 14:53:55 +01:00
reger	73ba5d8ef7	adjust fieldtype and description of field httpstatus_redirect_s in CollectionSchema - the field is not used (delete candidate)	2014-12-26 18:21:35 +01:00
Michael Peter Christen	eb78388a98	changed prefer strategy for http unique in such a way that http is preferred over https. While this is a bad idea from the standpoint of security it is more common applicable for environments where http and https mix and for some domains https is not available. Then the double-check is possible even if no postprocessing is performed.	2014-12-21 19:17:06 +01:00
Michael Peter Christen	aaf7d4775a	Merge branch 'master' of ssh://git@gitorious.org/yacy/rc1.git	2014-12-21 18:10:25 +01:00
Michael Peter Christen	8c3e5b7b6d	added experimental pdf splitting which enables YaCy to split pdfs during parsing into individual pages and add them all using different URLs. These constructed urls are generated from the source url with an appended page=<pagenumber> attribute to the url get/post properties. This will distinguish the different page entries. The search result list will then replace the post parameter with a url anchor # mark which causes that the original url is presented in the search result. These URLs can be opened directly on the correct page using pdf.js which is now built-in into firefox. That means: if you find a search hit on page 5 and click on the search result, firefox will open the pdf viewer and shows page 5.	2014-12-21 18:10:15 +01:00
reger	198102304b	refactor size() -> filesize() of URIMetadataNode (harmonize with ResultEntry and to not get confused with Collection.size())	2014-12-21 06:05:35 +01:00
Michael Peter Christen	d3e71ed070	fixes for searches when initialization of large autotagging libraries have not been finished	2014-12-19 17:38:58 +01:00
Michael Peter Christen	28683530cd	fixes to usage of no-cache: use and recognize also the no-store directive	2014-12-19 17:37:58 +01:00

... 3 4 5 6 7 ...

1497 Commits