Www.isocat.org ISOcat introduction 20 March 20121CLARIN-NL ISOcat workshop.

Slides:



Advertisements
Similar presentations
Dr. Leo Obrst MITRE Information Semantics Information Discovery & Understanding Command & Control Center February 6, 2014February 6, 2014February 6, 2014.
Advertisements

ISOcat Data Model: Workflow & Guidelines Marc Kemps-Snijders a, Sue Ellen Wright b, Menzo Windhouwer a a Max Planck Institute for Psycholinguistics, b.
ISOcat Data Category Registry Defining widely accepted linguistic concepts Menzo Windhouwer 1CLARIN-NL MD tutorial, September 2009.
Principles of ISOcat, a Data Category Registry Marc Kemps-Snijders a, Menzo Windhouwer a, Sue Ellen Wright b a Max Planck Institute for.
ISOcat introduction 19 June 20121CLARIN-NL ISOcat workshop.
Data Category specifications 19 June 20121CLARIN-NL 2012 ISOcat tutorial.
CLARIN-NL/VL procedure 20 June 20131CLARIN-NL ISOcat workshop.
11 CLARIN? ISOCAT! Ineke Schuurman ISOcat content coördinator CLARIN-NL Amsterdam
The Language Archive – Max Planck Institute for Psycholinguistics Nijmegen, The Netherlands Metadata Component Framework Possible Standardization Work.
MLIF: A Metamodel to Represent and Exchange Multilingual Textual Information ISO TC37 SC4 WG Samuel Cruz-Lara, Gil Francopoulo, Laurent Romary,
The current state of Metadata - as far as we understand it - Peter Wittenburg The Language Archive - Max Planck Institute CLARIN Research Infrastructure.
ISOcat: known issues 10 May /20111CLARIN-NL ISOcat workshop.
Data Category specifications 20 March 20121CLARIN-NL ISOcat workshop.
CLARIN-NL: Dealing with ISOcat Ineke Schuurman. ISOcat and CLARIN Projects call 1 CLARIN-NL Joint Flemish/Dutch pilot Whenever relevant, elements are.
CLARIN for Linguists Introduction Jan Odijk LOT Summerschool Nijmegen,
Tutorial for SC 32/WG 1 e-Business Standards Prepared for: SC Kunming Plenary Meeting Wenfeng Sun, Convenor ISO/IEC JTC1 SC32 WG1 (eBusiness)
9 th Open Forum on Metadata Registries Harmonization of Terminology, Ontology and Metadata 20th – 22nd March, 2006, Kobe Japan. Commonalities and Differences.
CLARIN-NL Second Open Call Jan Odijk CLARIN-NL Call 2 Info-session Amsterdam, 26 Aug 2010.
Agenda CMDI Workshop 9.15 Welcome 9.30 Introduction to metadata and the CLARIN Metadata Infrastructure (CMDI) 10.15Coffee 10.30Use of ISOCat within CMDI.
CLARIN-NL ISOcat workshop 2011 part 2 Ineke Schuurman Menzo Windhouwer.
The ISO-DCR 17 January /20111CMDI tutorial Marc Kemps-Snijders a, Menzo Windhouwer b, Sue Ellen Wright c a Meertens Institute, b MPI for.
Standards for language resources the ISO/TC 37(/SC 4) perspective
ISOcat demo and providing RELcat input Menzo Windhouwer The Language Archive tla.mpi.nl Data Archiving and Networked Solutions
LIRICS mid-term review 1 LIRICS WP3: Morpho-syntactic and syntactic annotations Thierry Declerck DFKI-LT - Saarbrücken 23rd May 2006.
Trends in Concept Modelling Turning Issues into Solutions How to Discipline a Cat Sue Ellen Wright, Kent State University.
INF 384 C, Spring 2009 Ontologies Knowledge representation to support computer reasoning.
CLARIN-NL Call 3 ISOcat follow-up 10/10/20121CLARIN-NL ISOcat Call 3 follow-up.
Content of the Data Category Registry 10 May /20111CLARIN-NL ISOcat workshop.
LIRICS Mid-term Review 1 LIRICS WP2 – NLP Lexica Monica Monachini CNR-ILC - Pisa 23rd May 2006.
CLARIN Metadata Infrastructure Component Metadata and intermediate solutions Daan Broeder Claus Zinn Dieter van Uytvanck - Max-Planck Institute for Psycholinguistics.
Nancy Lawler U.S. Department of Defense ISO/IEC Part 2: Classification Schemes Metadata Registries — Part 2: Classification Schemes The revision.
ISOcat: known issues 20 June 20131CLARIN-NL ISOcat workshop.
Report on the ISOcat project Marc Kemps-Snijders Menzo Windhouwer Peter Wittenburg Sue Ellen Wright January 8,
24 Jan 2005 Kick off meeting (Luxembourg) 1 LIRICS Linguistic Infrastructure for Interoperable Resources and Systems ►Kick off meeting presentation ►Proposal.
24 Jan 2005 Kick off meeting (Luxembourg) 1 LIRICS Linguistic Infrastructure for Interoperable Resources and Systems ►Kick off meeting presentation ►Proposal.
CLARIN-NL Call 4 ISOcat follow-up 2/10/20131CLARIN-NL Call 4 ISOcat follow-up.
ISOcat introduction 20 June 20131CLARIN-NL ISOcat workshop.
CLARIN-NL ISOcat workshop 2012 part 2 ( ) Ineke Schuurman Menzo Windhouwer.
ISOcat: known issues 19 June 20121CLARIN-NL ISOcat workshop.
11 CMDI/ISOcat And Semantic Operability Ineke Schuurman ISOcat content coördinator CLARIN-NL Menzo Windhouwer ISOcat system administrator Utrecht
The Language Archive – Max Planck Institute for Psycholinguistics Nijmegen, The Netherlands NP CMDI-1 Metadata Component Framework New Standardization.
A Data Category Registry- and Component- based Metadata Framework Daan Broeder et al. Max-Planck Institute for Psycholinguistics LREC 2010.
ISOcat: How to create a DC (including “do’s and don’ts”) 19 June 20121CLARIN-NL ISOcat tutorial.
Beyond ISOcat 20 June 2013CLARIN-NL ISOcat tutorial1.
Agenda CMDI Tutorial 9.30 Welcome & Coffee Introduction to metadata and the CLARIN Metadata Infrastructure (CMDI) 10.30CMDI & ISO-DCR 10.50The CMDI.
ISO TC 37/CLARIN SEMANTIC DATA REGISTRY WORKSHOP UTRECHT, DECEMBER ISOcat: Metadata Registry SUE ELLEN WRIGHT DECEMBER 2013.
Issues in Ontology-based Information integration By Zhan Cui, Dean Jones and Paul O’Brien.
CLARIN Concept Registry: the new semantic registry Ineke Schuurman, Menzo Windhouwer, Oddrun Ohren, Daniel Zeman
Towards a roadmap for standardization in language technology Laurent Romary & Nancy Ide Loria-INRIA — Vassar College.
ISOcat status
CLARIN Requirements for a Semantic Registry Daan Broeder The Language Archive – MPI Ineke Schuurman CLARIN-NL/VL – KU Leuven & Utrecht.
Extending the MDR for Semantic Web November 20, 2008 SC32/WG32 Interim Meeting Vilamoura, Portugal - Procedure for the Specification of Web Ontology -
1 ISOCAT Proposed solutions for Problems encountered in DUELME-LMF Jan Odijk Nijmegen 21 Sep 2010.
1 CLARIN? ISOCAT! Ineke Schuurman Hilversum,
Creating & Testing CLARIN Metadata Components A CLARIN-NL project Folkert de Vriend Meertens Institute, Amsterdam 18/05/2010.
The FDES revision process: progress so far, state of the art, the way forward United Nations Statistics Division.
SemAF – Basics: Semantic annotation framework Harry Bunt Tilburg University isa -6 Joint ISO - ACL/SIGSEM workshop Oxford, January 2011 TC 37/SC.
Annotation by category – ELAN and ISO DCR Han Slöetjes, Peter Wittenburg Max-Planck-Institute for Psycholinguistics LREC,
Be.wi-ol.de User-friendly ontology design Nikolai Dahlem Universität Oldenburg.
ISO TC 37/CLARIN DISCUSSION UTRECHT, DECEMBER 9/ Thinning Down a Bloated Cat SUE ELLEN WRIGHT DECEMBER 2013.
ISOcat tutorial DCR data model and guidelines. Simple and complex DCs Simple Data CategoryComplex Data CategoryConceptual Domain Data CategoryDescription.
A Data Category Registry- and Component- based Metadata Framework Daan Broeder et al. Max-Planck Institute for Psycholinguistics LREC 2010.
ISOcat: How to create a DC (including “do’s and don’ts”) 20 June 20131CLARIN-NL ISOcat tutorial.
TDS-Curator DANS MPI for Psycholinguistics Utrecht Institute of Linguistics OTS languagelink.let.uu.nl/tds/ 9/21/20101CLARIN-NL - Call 1 - ISOcat status.
Group work and standardization features in ISOcat Menzo Windhouwer 8/14/20101Standardizing Data Categories in ISOcat - Implementing Group.
Linking to Linguistic Data Categories in ISOcat Menzo Windhouwer a, Sue Ellen Wright b a The Language Archive - MPI for Psycholinguistics,
ISOcat introduction 10 May /20111CLARIN-NL ISOcat workshop.
WP4 Models and Contents Quality Assessment
Marc Kemps-Snijders Menzo Windhouwer Sue Ellen Wright
2. An overview of SDMX (What is SDMX? Part I)
Presentation transcript:

ISOcat introduction 20 March 20121CLARIN-NL ISOcat workshop

ISOcat: a Data Category Registry An implementation of ISO 12620:2009 – Terminology and other content and language resources — Specification of data categories and management of a Data Category Registry for language resources Successor to ISO 12620:1999 which contained a hardcoded list of Data Categories A data category – is the result of the specification of a given data field – an elementary descriptor in a linguistic structure or an annotation scheme 20 March 20122CLARIN-NL ISOcat workshop

What is a Data Category? The result of the specification of a given data field – A data category is an elementary descriptor in a linguistic structure or an annotation scheme. Specification consists of 3 main parts: – Administrative part Administration and identification – Descriptive part Documentation in various working languages – Linguistic part Conceptual domain(s for various object languages) 20 March 2012CLARIN-NL ISOcat workshop3

Data Category example Data category: /Grammatical gender/ – Administrative part: Identifier: grammaticalGender PID: – Descriptive part: English definition: Category based on (depending on languages) the natural distinction between sex and formal criteria. French definition: Catégorie fondée (selon la langue) sur la distinction naturelle entre les sexes ou d'autres critères formels. – Linguistic part: Morposyntax conceptual domain: /male/, /feminine/, /neuter/ French conceptual domain: /male/, /feminine/ 20 March 2012CLARIN-NL ISOcat workshop4

Data Category types 20 March 2012CLARIN-NL ISOcat workshop5 writtenForm string open grammaticalGender string neuter masculine feminine closed simple: string constrained complex:

Data Category types 20 March 2012CLARIN-NL ISOcat workshop6 language alphabet writtenForm japanese ipa lexicon entry lemma container:

20 March 2012CLARIN-NL ISOcat workshop7 Data Category relationships Value domain membership Subsumption relationships between simple data categories (legacy) Relationships between complex/container data categories are not stored in the DCR partOfSpeech string pronoun personal pronoun

20 March 2012CLARIN-NL ISOcat workshop8 No ontological relationships? Rationale: – Relation types and modeling strategies for a given data category may differ from application to application; – Motivation to agree on relation and modeling strategies will be stronger at individual application level; – Integration of multiple relation structures in DCR itself could lead to endless ontological clutter. Solution under development: RELcat a Relation Registry

How can you use Data Categories? Lexicon Lexical Entry FormSense 0..* 1..* Word Form Lemma LanguageBWOgenders grammaticalGenderwordOrder A (schema for a) lexicon A (schema for a) typological database 20 March 20129CLARIN-NL ISOcat workshop partOfSpeech writtenForm grammaticalGender lexicalType lemma wordForm lexicalEntry lexicon Shared semantics!

20 March 2012CLARIN-NL ISOcat workshop10 What is a Data Category Registry? A (coherent) set of Data Categories, in our case for linguistic resources A system to manage this set: – Create and edit Data Categories – Share Data Categories, e.g., resolve PID references – Standardize Data Categories Grass roots approach

Standardization Submission group Data Category Registry Board Validation Thematic Domain Group Evaluation Stewardship group Decision Group rejected Publication 20 March CLARIN-NL ISOcat workshop

20 March 2012CLARIN-NL ISOcat workshop12 Thematic Domain Groups TDG 1: Metadata TDG 2: Morphosyntax TDG 3: Semantic Content Representation TDG 4: Syntax TDG 5: Machine Readable Dictionary TDG 6: Language Resource Ontology TDG 7: Lexicography TDG 8: Language Codes TDG 9: Terminology TDG 11: Multilingual Information Management TDG 12: Lexical Resources TDG 13: Lexical Semantics TDG 14: Source Identification TDGs are the owner and guardians of a coherent subset of the DCR TDGs own one or more profiles Each TDG has a chair A number of judges (assigned by SC P members) A number of expert members (up to 50%) TDGs are constituted at the TC37/SC plenary New TDGs need to be proposed by a SC 1.Translation 2.Sign language 3.Audio

20 March 2012CLARIN-NL ISOcat workshop13 How can you use a Data Category Registry? You can: – Find Data Categories relevant for your resources and embed references to them so the semantics of (parts of) your resources are made explicit This can be supported by tools you use, e.g., ELAN, LEXUS and the CMDI Component Editor directly interact with ISOcat – Interact with Data Category owners to improve (the coverage of) their Data Categories – Create (together with others) new Data Categories and/or selections needed for your resources and share those – Submit (your) Data Categories for standardization – Free of charge – Grass roots approach

ISOcat and CLARIN(-NL/VL): general remarks 20 March CLARIN-NL ISOcat workshop

Importance of ISOcat Collaboration – Human, machine, language x, language y Essential in CLARIN, but … Impossible when we don’t know (exactly) what we are talking about! -Transitive verb – transitief werkwoord -Transitief werkwoord – overgankelijk werkwoord 20 March 2012CLARIN-NL ISOcat workshop15

Importance of ISOcat ISOcat: – Provides us with a framework to make such things clear (is X the same as Y, does A use it the same way) – At least, that is the intention, ISOcat still being ‘under construction’ Today’s sessions: – How to work with ISOcat – Which other “cats” do we have at the moment – The future … 20 March 2012CLARIN-NL ISOcat workshop16

CLARIN-NL (and VL) and ISOcat There are some 40 projects dealing with ISOcat in some sense (sometimes ‘only’ metadata) – 35 Netherlands – 3 Flanders – 1 NL/VL pilot – Of course, that is not the main focus of these projects, but still… – A lot of ISOcat work needs to be done! 20 March 2012CLARIN-NL ISOcat workshop17

CLARIN-NL (and VL) and ISOcat At least of TTNWW (the pilot) one of the explicit goals is to signal problems and to try to remedy them (for our own good, and that of CLARIN as a whole) In that respect, we do have some ‘success’ – Several larger and smaller issues are already being remedied At l 20 March 2012CLARIN-NL ISOcat workshop18

CLARIN-NL (and VL) and ISOcat Many (Dutch) projects working on ISOcat issues, plus those of other national CLARINs same concepts ? same problems ?  very likely 20 March 2012CLARIN-NL ISOcat workshop19

Collaboration necessary National (Dutch) level Coordinated effort Shared workspace under ‘shared’ USE IT Plus discussion platform Report problems to me (Ineke) International level We will try to collaborate with them as well 20 March 2012CLARIN-NL ISOcat workshop20

Thanks ! 20 March 2012CLARIN-NL ISOcat workshop21