My Account Log in

2 options

Spanish gigaword / [Author, Dave Graff].

LIBRA -
Loading location information...

Available from offsite location This item is stored in our repository but can be checked out.

Log in to request item
LIBRA -
Loading location information...

Available from offsite location This item is stored in our repository but can be checked out.

Log in to request item
Format:
Datafile
Contributor:
Graff, David Andrew, 1962-
Linguistic Data Consortium.
Agence France-presse.
Associated Press.
Xin hua tong xun she.
Language:
English
Spanish
Subjects (All):
Spanish language--Data processing--Databases.
Spanish language.
Information retrieval.
Natural language processing (Computer science).
Foreign news--Databases.
Foreign news.
Spanish language--Data processing.
Genre:
Databases.
Academic theses.
Physical Description:
1 DVD-ROM ; 4 3/4 in.
4 3/4 in.
Edition:
First edition.
Other Title:
Spanish gigaword first edition
Place of Publication:
[Philadelphia, Pa.] : Linguistic Data Consortium, [2006]
Language Note:
Spanish and English.
System Details:
digital
optical
data file
Summary:
"This file contains documentation on the Spanish Gigaword First Edition, Linguistic Data Consortium (LDC) catalog number LDC2006T12 and isbn 1-58563-393-3. The Spanish Gigaword Corpus is a comprehensive archive of newswire text data that has been acquired over several years by the Linguistic Data Consortium (LDC) at the University of Pennsylvania. This is the first edition of the Spanish Gigaword Corpus, though some of the data included here has been released previously in other LDC corpora. The three distinct international sources of Spanish newswire in this edition, and the time spans of collection covered for each, are as follows: Agence France-Presse, Spanish Service (afp_spa) May 1994 - Dec 2005 ; Associated Press Worldstream, Spanish (apw_spa) Nov 1993 - Dec 2005 ; Xinhua News Agency, Spanish Service (xin_spa) Sep 2001 - Dec 2005. The seven-letter codes in the parentheses above include the three-character source name abbreviations and the three-character language code ("spa") separated by an underscore ("_") character. The three-letter language code conforms to LDC's new internal convention based on the new ISO 639-3 standard. The seven-letter codes are used in both the directory names where the data files are found, and in the prefix that appears at the beginning of every data file name. It is also used (in all UPPER CASE) as the initial portion of the DOC "id" strings that uniquely identify each news story."--index.html.
Notes:
"LDC2006T12."
Title from index.html on DVD-ROM.
ISBN:
1585633933
9781585633937
OCLC:
72463518

The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.

Find

Home Release notes

My Account

Shelf Request an item Bookmarks Fines and fees Settings

Guides

Using the Find catalog Using Articles+ Using your account