Finding Efficient Linguistic Feature Set for Authorship Verification

dc.contributor.authorRanatunga, R.V.S.P.K.
dc.date.accessioned2022-04-06T10:11:10Z
dc.date.available2022-04-06T10:11:10Z
dc.date.issued2013
dc.description.abstractAuthorship verification rely on identification of a given document to verify whether it is written by a particular author or not. Internally, analyzing the document itself with respect to variations in writing style of the author and identification of the author‟s own idiolect is the main context of the authorship verification. Mainly, the detection performance depends on the used feature set for clustering the document. Linguistic features and stylistic features have been utilized for author identification according to the writing style of a particular author. Disclosing the shallow changes of the author‟s writing style is the major problem which should be addressed in the domain of authorship verification. It motivates the computer science researchers to do research on authorship verification in the field of computer forensics and this research also focuses on this problem. The contributions from the proposed research are two folded: Former is introducing a new feature extracting method with Natural Language Processing (NLP) and latter is proposing a novel and more efficient linguistic feature set for verification of the author of the given document. Experiments were carried out on a corpus composed of freely downloadable genuine 19th century English text. Each word segment obtained from the corpus is subjected to feature extraction and 49 stylistic features are used for clustering the text. Other than the standard stylistic features, 19 linguistic features are used as new feature set for the experiments. Generated parse trees by the Stanford Parser are utilized for extracting these linguistic features. Self organizing maps have been used as the classifier to cluster the documents. Proper word segmentation is also introduced in this work which helps us to demonstrate that the proposed strategy can produce promising results. Finally, it is realized that more accurate classification is generated by the proposed strategy with the extracted linguistic feature set.en_US
dc.identifier.citationRanatunga, R.V.S.P.K.(2013).Finding Efficient Linguistic Feature Set for Authorship Verification, Journal of Computer Science Vol. 1, No. 1 (2013) 35-43en_US
dc.identifier.doihttps://doi.org/10.31357/jcs.v1i1.1616en_US
dc.identifier.urihttp://dr.lib.sjp.ac.lk/handle/123456789/11014
dc.language.isoenen_US
dc.publisherDepartment of Computer Science Faculty of Applied Sciences University of Sri Jayewardenepuraen_US
dc.subjectAuthorship Verification, Style Markers, Natural Language Processing, Self Organizing Mapsen_US
dc.titleFinding Efficient Linguistic Feature Set for Authorship Verificationen_US
dc.typeArticleen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
admin,+linguistic-proof.pdf
Size:
389.52 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: