Normalizace dat ve fulltextovém systému
The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservat...
Main Author: | |
---|---|
Other Authors: | |
Format: | Dissertation |
Language: | Slovak |
Published: |
2010
|
Online Access: | http://www.nusl.cz/ntk/nusl-286730 |
id |
ndltd-nusl.cz-oai-invenio.nusl.cz-286730 |
---|---|
record_format |
oai_dc |
spelling |
ndltd-nusl.cz-oai-invenio.nusl.cz-2867302017-06-27T04:41:16Z Normalizace dat ve fulltextovém systému Normalisation of data in fulltext system Knap, Tomáš Kapusta, Matúš Lánský, Jan The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservation of the structure and conservation of the text content as much as possible. Program verifies if namespaces are correct and allows replacement of special HTML entities by appropriate Unicode characters. The output file has to be correctly processed by standard XML parsers. Secondary objective is to investigate and implement suitable method to identify language of the resulting output file as a whole, the content of individual elements or plain text file. 2010 info:eu-repo/semantics/masterThesis http://www.nusl.cz/ntk/nusl-286730 slo info:eu-repo/semantics/restrictedAccess |
collection |
NDLTD |
language |
Slovak |
format |
Dissertation |
sources |
NDLTD |
description |
The purpose of this thesis is to design and implement Java application to process data from full text system to well-formed XML data according to XML 1.0 specification. Input data are stored in XML files containing any kind of not well formed XML, typically HTML content. Major criteria are conservation of the structure and conservation of the text content as much as possible. Program verifies if namespaces are correct and allows replacement of special HTML entities by appropriate Unicode characters. The output file has to be correctly processed by standard XML parsers. Secondary objective is to investigate and implement suitable method to identify language of the resulting output file as a whole, the content of individual elements or plain text file. |
author2 |
Knap, Tomáš |
author_facet |
Knap, Tomáš Kapusta, Matúš |
author |
Kapusta, Matúš |
spellingShingle |
Kapusta, Matúš Normalizace dat ve fulltextovém systému |
author_sort |
Kapusta, Matúš |
title |
Normalizace dat ve fulltextovém systému |
title_short |
Normalizace dat ve fulltextovém systému |
title_full |
Normalizace dat ve fulltextovém systému |
title_fullStr |
Normalizace dat ve fulltextovém systému |
title_full_unstemmed |
Normalizace dat ve fulltextovém systému |
title_sort |
normalizace dat ve fulltextovém systému |
publishDate |
2010 |
url |
http://www.nusl.cz/ntk/nusl-286730 |
work_keys_str_mv |
AT kapustamatus normalizacedatvefulltextovemsystemu AT kapustamatus normalisationofdatainfulltextsystem |
_version_ |
1718469392221601792 |