Normalizace dat ve fulltextovém systému

The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservat...

Full description

Bibliographic Details
Main Author:	Kapusta, Matúš
Other Authors:	Knap, Tomáš
Format:	Dissertation
Language:	Slovak
Published:	2010
Online Access:	http://www.nusl.cz/ntk/nusl-286730

id	ndltd-nusl.cz-oai-invenio.nusl.cz-286730
record_format	oai_dc
spelling	ndltd-nusl.cz-oai-invenio.nusl.cz-2867302017-06-27T04:41:16Z Normalizace dat ve fulltextovém systému Normalisation of data in fulltext system Knap, Tomáš Kapusta, Matúš Lánský, Jan The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservation of the structure and conservation of the text content as much as possible. Program verifies if namespaces are correct and allows replacement of special HTML entities by appropriate Unicode characters. The output file has to be correctly processed by standard XML parsers. Secondary objective is to investigate and implement suitable method to identify language of the resulting output file as a whole, the content of individual elements or plain text file. 2010 info:eu-repo/semantics/masterThesis http://www.nusl.cz/ntk/nusl-286730 slo info:eu-repo/semantics/restrictedAccess
collection	NDLTD
language	Slovak
format	Dissertation
sources	NDLTD
description	The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservation of the structure and conservation of the text content as much as possible. Program verifies if namespaces are correct and allows replacement of special HTML entities by appropriate Unicode characters. The output file has to be correctly processed by standard XML parsers. Secondary objective is to investigate and implement suitable method to identify language of the resulting output file as a whole, the content of individual elements or plain text file.
author2	Knap, Tomáš
author_facet	Knap, Tomáš Kapusta, Matúš
author	Kapusta, Matúš
spellingShingle	Kapusta, Matúš Normalizace dat ve fulltextovém systému
author_sort	Kapusta, Matúš
title	Normalizace dat ve fulltextovém systému
title_short	Normalizace dat ve fulltextovém systému
title_full	Normalizace dat ve fulltextovém systému
title_fullStr	Normalizace dat ve fulltextovém systému
title_full_unstemmed	Normalizace dat ve fulltextovém systému
title_sort	normalizace dat ve fulltextovém systému
publishDate	2010
url	http://www.nusl.cz/ntk/nusl-286730
work_keys_str_mv	AT kapustamatus normalizacedatvefulltextovemsystemu AT kapustamatus normalisationofdatainfulltextsystem
_version_	1718469392221601792

Normalizace dat ve fulltextovém systému

Similar Items