Normalizace dat ve fulltextovém systému

The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservat...

Full description

Bibliographic Details
Main Author: Kapusta, Matúš
Other Authors: Knap, Tomáš
Format: Dissertation
Language:Slovak
Published: 2010
Online Access:http://www.nusl.cz/ntk/nusl-286730
id ndltd-nusl.cz-oai-invenio.nusl.cz-286730
record_format oai_dc
spelling ndltd-nusl.cz-oai-invenio.nusl.cz-2867302017-06-27T04:41:16Z Normalizace dat ve fulltextovém systému Normalisation of data in fulltext system Knap, Tomáš Kapusta, Matúš Lánský, Jan The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservation of the structure and conservation of the text content as much as possible. Program verifies if namespaces are correct and allows replacement of special HTML entities by appropriate Unicode characters. The output file has to be correctly processed by standard XML parsers. Secondary objective is to investigate and implement suitable method to identify language of the resulting output file as a whole, the content of individual elements or plain text file. 2010 info:eu-repo/semantics/masterThesis http://www.nusl.cz/ntk/nusl-286730 slo info:eu-repo/semantics/restrictedAccess
collection NDLTD
language Slovak
format Dissertation
sources NDLTD
description The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservation of the structure and conservation of the text content as much as possible. Program verifies if namespaces are correct and allows replacement of special HTML entities by appropriate Unicode characters. The output file has to be correctly processed by standard XML parsers. Secondary objective is to investigate and implement suitable method to identify language of the resulting output file as a whole, the content of individual elements or plain text file.
author2 Knap, Tomáš
author_facet Knap, Tomáš
Kapusta, Matúš
author Kapusta, Matúš
spellingShingle Kapusta, Matúš
Normalizace dat ve fulltextovém systému
author_sort Kapusta, Matúš
title Normalizace dat ve fulltextovém systému
title_short Normalizace dat ve fulltextovém systému
title_full Normalizace dat ve fulltextovém systému
title_fullStr Normalizace dat ve fulltextovém systému
title_full_unstemmed Normalizace dat ve fulltextovém systému
title_sort normalizace dat ve fulltextovém systému
publishDate 2010
url http://www.nusl.cz/ntk/nusl-286730
work_keys_str_mv AT kapustamatus normalizacedatvefulltextovemsystemu
AT kapustamatus normalisationofdatainfulltextsystem
_version_ 1718469392221601792