Mostrando entradas con la etiqueta Biomedicine. Mostrar todas las entradas
Mostrando entradas con la etiqueta Biomedicine. Mostrar todas las entradas

23.11.11

Datasets, databases and resources: MSH WSD, BioNOT, Gazetiki, DBpedia Spotlight, Google BigQuery, Common Crawl

Some datasets and resources I have recently found (although they may be old):

  • MSH WSD: a data set for Word Sense Disambiguation WSD based on a method that can be used to automatically develop a WSD test collection using the Unified Medical Language System (UMLS) Metathesaurus and the manual MeSH indexing of MEDLINE.
  • BioNOT: a searchable database of negated biomedical sentences. The database consists of more than 32 million negated sentences at PubMed.
  • Gazetiki: a geographical database that contains 8323702 geographical names coming from Geonames and from different Web sources, with the latter representing over 1 million items, with the addition of a popularity score which was calculated based on the usage of a place name in a geotagged dataset.
  • DBpedia Spotlight: a tool for automatically annotating mentions of DBpedia resources in text, providing a solution for linking unstructured information sources to the Linked Open Data cloud through DBpedia. DBpedia Spotlight performs named entity extraction, including entity detection and Name Resolution.
  • Google BigQuery Service: a SQL-like tool for analyzing massive datasets, as a web service that enables you to do interactive analysis of massively large datasets-up to billions of rows.
  • Common Crawl: a freely accessible index of 5 billion web pages, their page rank, their link graphs and other metadata, hosted on Amazon EC2, was announced today by the Common Crawl Foundation.

14.5.10

CFP: Workshop on Language Technology applied to biomedical documents

Workshop on Language Technology applied to biomedical documents
SEPLN 2010 satellite workshop

In the last decade, language technology has received an increasing interest as suitable solution to retrieval and analyse the huge volume of published documents in biological domain. Recently, medical domain also benefit from the application of such technology. The workshop is intended to provide a forum for discussing the latest advances of language technology applied to biological and medical domains. The workshop aims to provide a broad view on the shortcomings of current existing techniques, tools, or resources as well as emergent applications concerning accessing scientific publications and health general interest documents with special attention to non English documents.

Authors are invited to submit original papers addressing any of the following key topics but not limited to:

  • Text mining from clinical documents
  • Integration of biomedical resources in specific applications
  • Woks on minor languages and/or different from English
  • Real world applications (IR systems for medical and scientific specialists, medical education, health knowledge organization, ERH...)
  • Evaluation methodologies
  • Biomedical corpus development
  • Information Retrieval in health domain
  • Classification of clinical and biological documents (for instance, ICD-10)
  • Biomedical Named Entity Recognition and Concept Identification
  • Information Extraction from biological and clinical documents
  • Anonymisation of clinical texts
  • Information fusion: integrating data from heterogeneous biomedical sources, connecting resources
  • Methods of creating, reviewing and editing scientific content
  • Summarization of electronic patient records, medical reports, scientific articles, etc.
  • Creation of biomedical annotated corpora
  • Creation and evaluation of linguistic tools for biomedical domain in different languages
  • Evaluation methodologies in biomedical domain: system-oriented and user-oriented evaluations

Important dates

  • Paper submission: June 17th
  • Notification of acceptance for papers: July 2nd
  • Final Camera Ready paper due: July 15th
  • Worshop day: September, 6 or 7th 2010

16.2.10

CFP: Intelligent Methods for Protecting Privacy and Confidentiality in Data

Intelligent Methods for Protecting Privacy and Confidentiality in Data
May 30th, 2010, Ottawa, Canada

Submission deadline: March 30th

With the increasing adoption of electronic medical/health records and the rising use of electronic data capture tools in clinical research, large electronic repositories of personal health information (PHI) are being built up. At the same time, large medical data breaches are becoming common. Data breaches may be caused by errors committed by insiders at the data custodian sites, or by malicious insiders. Data
breaches can also be caused by outsiders breaking into the data repositories. These data breaches represent legal and financial liabilities for the data custodians, and erode public trust in the ability of data custodians to manage their PHI.

An area that has grown in importance to manage the risks from breaches is data leak prevention (DLP). DLP technologies monitor communications or networks to detect PHI leaks. When a leak is detected the affected individual or organization is notified, at which point they can take remedial action. DLP can prevent a PHI leak or detect it after it happens. For example, if DLP is deployed to monitor email then a PHI alert can be generated before the email is sent. If DLP is used to monitor PHI leaks on the Internet (e.g., on peer-to-peer file sharing networks or on web sites), then the alerts pertain to leaks that have already occured, at which point the affected individual or data custodian can attempt to contain the damage and stop further leaks.

Computational AI is a key enabling technology for next-generation DLP technologies. This workshop aims to bring together researchers working on computational tools for DLP.

Topics of interest include, but are not limited to:

  • reviews: reviews of DLP systems and methods; and reviews of PHI leaks that are occuring.
  • methods: detection of personally identifying information in text; detection of health information in different types of text (e.g., professionally written vs. lay person generated); and re- identification risk assessment;
  • applications: monitoring the web and peer-to-peer file sharing networks for PHI leaks; detection of PHI in email or other communications; and tools for dealing with PHI leaks in an automated way (e.g., de-identification).
  • evaluation: empirical evaluation of deployed systems; theoretical methods of risk assessment; and new methods for evaluating such systems.

Workshop Format

The workshop invites position papers describing original work in theory and applications of intelligent methods to the problem of DLP. Position papers will be reviewed by the Program Committee members according to their originality, technical merit and clarity of presentation. Each accepted paper will be allocated a maximum of 5 pages in the workshop proceedings. At least one author for each accepted paper is expected to attend the workshop.

The workshop is planned to be interactive with discussions on the current state and future developments in the area of DLP for PHI. All of the workshop attendees will co-author a final report on DLP for PHI after the workshop and submit that to a journal.

Location

The workshop is being held in conjunction with the Canadian AI 2010 conference.

11.2.10

CFP: Workshop on Intelligent Methods for Protecting Privacy and Confidentiality in Data

Workshop on Intelligent Methods for Protecting Privacy and Confidentiality in Data
May 30th, 2010, Ottawa

With the increasing adoption of electronic medical/health records and the rising use of electroinc data capture tools in clinical research, large electronic repositories of personal health information (PHI) are being built up. At the same time, large medical data breaches are becoming common. Data breaches may be caused by errors committed by insiders at the data custodian sites, or by malicious insiders. Data breaches can also be caused by outsiders breaking into the data repositories. These data breaches represent legal and financial liabilities for the data custodians, and erode public trust in the ability of data custodians to manage their PHI.

An area that has grown in importance to manage the risks from breaches is data leak prevention (DLP). DLP technologies monitor communications or networks to detect PHI leaks. When a leak is detected the affected individual or organization is notified, at which point they can take remedial action. DLP can prevent a PHI leak or detect it after it happens. For example, if DLP is deployed to monitor email then a PHI alert can be generated before the email is sent. If DLP is used to monitor PHI leaks on the Internet (e.g., on peer-to-peer file sharing networks or on web sites), then the alerts pertain to leaks that have already occured, at which point the affected individual or data custodian can attempt to contain the damage and stop further leaks.

Computational AI is a key enabling technology for next-generation DLP technologies. This workshop aims to bring together researchers working on computational tools for DLP.

Topics of interest include, but are not limited to:

  • reviews
    • reviews of DLP systems and methods; and
    • reviews of PHI leaks that are occuring.
  • methods
    • detection of personally identifying information in text;
    • detection of health information in different types of text (e.g., professionally written vs. lay person generated); and
    • re-identification risk assessment;
  • applications
    • monitoring the web and peer-to-peer file sharing networks for PHI leaks;
    • detection of PHI in email or other communications; and
    • tools for dealing with PHI leaks in an automated way (e.g., de-identification).
  • evaluation
    • empirical evaluation of deployed systems;
    • theoretical methods of risk assessment; and
    • new methods for evaluating such systems.

Workshop Format

The workshop invites position papers describing original work in theory and applications of intelligent methods to the problem of DLP. Position papers will be reviewed by the Program Committee members according to their originality, technical merit and clarity of presentation. Each accepted paper will be allocated a maximum of 5 pages in the workshop proceedings. At least one author for each accepted paper is expected to attend the workshop.

The workshop is planned to be interactive with discussions on the current state and future developments in the area of DLP for PHI. All of the workshop attendees will co-author a final report on DLP for PHI after the workshop and submit that to a journal.

Location

The workshop is being held in conjunction with the Canadian AI 2010 conference. Location and registration information is available at its web page.

Important Dates

  • Full paper submission: March 30, 2010
  • Notification of acceptance: April 15, 2010
  • Camera-ready submission: May 1, 2010

2.9.09

Text Mining Hands-on course and training seminar at Cambridge, U.K., 5th-6th October 2009

The European Bioinformatics Institute (EBI) and the National Centre for Text Mining (University of Manchester) are organising a joint training event at the EBI, on October 5th / 6th, 2009.

The purpose of this event is to teach basic techniques in information retrieval (IR) and information extraction (IE) in the biomedical domain and to give hands-on training on existing solutions provided by the two centres. This seminar will give you the opportunity to meet the experts behind the established solutions.

Intended audience: Biomedical researchers, biocurators, bioinformaticians, medical informaticians and any other researcher active in biomedical research.

You have to register till September 10th, 2009.

23.4.09

The Open Health Natural Language Processing Consortium

The goal of the Open Health Natural Language Processing Consortium is to establish an open source consortium to promote past and current development efforts and to encourage participation in advancing future efforts. The purpose of this consortium is to facilitate and encourage new annotator and pipeline development, exchange insights and collaborate on novel biomedical natural language processing systems and develop gold-standard corpora for development and testing. The Consortium promotes the open source UIMA framework and SDK as the basis for biomedical NLP systems. Applications created within UIMA consist of software components (referred to as annotators) and their associated configuration files and external resources. Within the framework, one can also create complete pipelines composed of a sequence of annotators and the data flow between them.

Via the BioNLP list.

Data resources for Biological / Chemical Natural Language Processing

In the list I am subscribed, there are announces of new data collections from time to time. These resources are extremely valuable for Bio-NLP research, we must disseminate them and strongly appreciate their builders work:

A number of tools are available at the Bio-NLP Resources page compiled by Martin Krallinger and his group.

14.4.09

Parsed MEDLINE(R) data download service

Jin-Dong Kim, from the Tsujii Laboratory, University of Tokyo, has announced the start of the Parsed MEDLINE(R) data download service.

This service provides an access to syntactically-parsed MEDLINE abstracts. The abstracts were parsed with a wide-coverage HPSG parser, the Enju parser (version 2.2). The original data is the 2009 baseline release of the MEDLINE database, which includes approximately 18 million records.

For detail, refer to the usage web page.

24.3.09

GENIA treebank corpus version 1.0

The GENIA treebank corpus version 1.0 is available from now at the GENIA project homepage: http://www-tsujii.is.s.u-tokyo.ac.jp/GENIA/

The corpus contains 1,999 PubMed abstracts with part-of-speech and syntactic tree annotation.

Seen through the BioNLP list.