<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://www.lotico.com/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Marco</id>
	<title>lotico - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://www.lotico.com/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Marco"/>
	<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php/Special:Contributions/Marco"/>
	<updated>2026-07-29T10:08:09Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.46.0</generator>
	<entry>
		<id>https://www.lotico.com/index.php?title=Semantic_Web_Jobs&amp;diff=6653</id>
		<title>Semantic Web Jobs</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Semantic_Web_Jobs&amp;diff=6653"/>
		<updated>2026-06-20T07:41:20Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Vice President- Ontologist JPMorgan Chase NYC or NJ */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;If you&#039;d like to hire a Semantic Web expert or if you&#039;re looking for a Semantic Web position,&lt;br /&gt;
please send me a short note or an HTTP link.&lt;br /&gt;
If you&#039;re posting a position, please be sure to note whether or not the job is in the NYC area.&lt;br /&gt;
If you&#039;re looking for a position, please indicate whether you&#039;re willing to consider locations besides the NYC area.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Vice President- Ontologist JPMorgan Chase===&lt;br /&gt;
&lt;br /&gt;
https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/job/210760263&lt;br /&gt;
&lt;br /&gt;
Location: NYC or NJ&lt;br /&gt;
&lt;br /&gt;
6/20/2026&lt;br /&gt;
&lt;br /&gt;
The Firmwide Chief Data Office is responsible for maximizing the value and impact of data globally, in a highly governed way. It consists of several teams focused on accelerating JPMorgan Chase’s data, analytics, and AI journey, including data strategy, data impact optimization, privacy, data governance, transformation, and talent. We are looking for a data executive to help shape the strategy for how we make data available to power everything from new product development to Artificial Intelligence models. This leader will join the team responsible for setting the firmwide data publishing strategy and driving the adoption of the strategy across the firm.&lt;br /&gt;
&lt;br /&gt;
As a Vice President-Ontologist within the JP Morgan Chase team, you will be instrumental in shaping our knowledge representation. Your role will involve utilizing ontologies and taxonomies to enhance data interoperability and management, preparing our data for AI applications. Your responsibilities will range from engaging with and educating domain experts, to assessing standard ontologies and developing our organization-wide ontology. Your work will traverse multiple domains, influencing departments like Data &amp;amp; Analytics, Product, and Tech.&lt;br /&gt;
&lt;br /&gt;
Job responsibilities:&lt;br /&gt;
&lt;br /&gt;
Development and adoption of ontologies to represent complex domains&lt;br /&gt;
&lt;br /&gt;
Evaluate industry standard ontologies for adoption across JPMC&lt;br /&gt;
&lt;br /&gt;
Work closely with stakeholders, subject matter experts, product owners, and engineers to understand their use cases, requirements, and dependencies, critically assessing proposed solutions&lt;br /&gt;
&lt;br /&gt;
Provide expert input into the JP Morgan Chase’s firmwide data strategy&lt;br /&gt;
&lt;br /&gt;
Communicate complex ideas effectively to collaborators using precise terminology and relatable examples, and ask clarifying questions to define core meanings.&lt;br /&gt;
&lt;br /&gt;
Mentor fellow ontologists to ensure alignment with accepted practices, standards, objectives, key results, and strategic initiatives.&lt;br /&gt;
&lt;br /&gt;
Keep abreast of emerging trends and advancements in ontology engineering, knowledge representation, and semantic technologies.&lt;br /&gt;
&lt;br /&gt;
Balance timeliness with quality under tight deadlines, managing multiple priorities and partners.&lt;br /&gt;
&lt;br /&gt;
Ensure end-to-end relevance to stakeholder needs, from gathering competency questions to achieving successful integrations.&lt;br /&gt;
&lt;br /&gt;
Required qualifications, capabilities, and skills:&lt;br /&gt;
&lt;br /&gt;
3+ years of experience developing and managing ontologies for real-world applications&lt;br /&gt;
&lt;br /&gt;
Expertise in Data and Financial service standards such as ISO 20022, OWL, RDF, SKOS, and SHACL.&lt;br /&gt;
&lt;br /&gt;
Experience with ontology and taxonomy development process and tools (e.g., Protégé, TopBraid Composer, PoolParty, etc.).&lt;br /&gt;
&lt;br /&gt;
Structured thinker and effective communicator with excellent written communication skills. Ability to crisply articulate complex technical concepts to senior audiences with poise and confidence.&lt;br /&gt;
&lt;br /&gt;
Preferred qualifications, capabilities, and skills:&lt;br /&gt;
&lt;br /&gt;
Master&#039;s or Ph.D. in a field focused on ontology engineering, knowledge representation, or semantic technologies, such as Information Science, Library Science, Philosophy, Linguistics, or Computer Science.&lt;br /&gt;
&lt;br /&gt;
Experience with Financial sector data standards and ontologies&lt;br /&gt;
&lt;br /&gt;
Experience with program management and collaborative development best practices&lt;br /&gt;
&lt;br /&gt;
Understanding of large-scale, distributed, end-to-end systems.&lt;br /&gt;
&lt;br /&gt;
Knowledge of Data Governance and Data Management&lt;br /&gt;
&lt;br /&gt;
Contributions to the Ontology community, such as papers, conference presentations, industry standard contributions, or open source contributions (e.g., Github repos)&lt;br /&gt;
&lt;br /&gt;
===Senior Lead Software Engineer-Ontology and RDF - JPMorgan Chase===&lt;br /&gt;
&lt;br /&gt;
Location:  GLASGOW, LANARKSHIRE, United Kingdom &lt;br /&gt;
&lt;br /&gt;
As a Lead Software Engineer at JPMorgan Chase within the Identity and Access Management, Corporate Sector, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.&lt;br /&gt;
&lt;br /&gt;
https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/requisitions/preview/210472446/?keyword=Ontology&lt;br /&gt;
&lt;br /&gt;
===Lead Software Engineer- Ontology and RDF - JPMorgan Chase===&lt;br /&gt;
&lt;br /&gt;
Location:  GLASGOW, LANARKSHIRE, United Kingdom &lt;br /&gt;
&lt;br /&gt;
As a Lead Software Engineer at JPMorgan Chase within the Identity and Access Management , Corporate Sector, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.&lt;br /&gt;
&lt;br /&gt;
https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/requisitions/preview/210472450/?keyword=Ontology&lt;br /&gt;
&lt;br /&gt;
===Ontologist IMDb Bristol===&lt;br /&gt;
&lt;br /&gt;
URL: Allocated &amp;lt;!-- https://www.amazon.jobs/en-gb/jobs/2026698/ontologist-imdb-content --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Location: Bristol, UK&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
DESCRIPTION&lt;br /&gt;
Job summary&lt;br /&gt;
With more than 400 million searchable data items — including 10 million movie, TV and entertainment titles, 11 million cast and crew members and 11 million images — IMDb is the world’s most popular and authoritative source for information on movies, TV shows and celebrities, and has a combined web and mobile audience of more than 200 million monthly visitors. The IMDb database is continually growing, thanks to a vast contributor community of entertainment professionals and companies, IMDb staff, individual contributors and other trusted sources. IMDb content is integrated into strategically important parts of Amazon and AWS businesses, including Amazon Fire TV, Alexa, and X-Ray on Prime Video. IMDb licenses information from its vast and authoritative database to third-party businesses, including film studios, television networks, streaming services and cable companies, as well as airlines, electronics manufacturers, non-profit organizations and software developers. Learn more at developer.imdb.com. Other IMDb products and services include: the IMDb website for desktop and mobile devices; apps for iOS and Android; a free streaming channel, IMDb TV; and IMDb original video series and podcasts. For entertainment industry professionals, IMDb provides IMDbPro and Box Office Mojo. IMDb is an Amazon company. For more information, visit imdb.com/press and follow @IMDb.&lt;br /&gt;
&lt;br /&gt;
IMDb is a group of entertainment enthusiasts – and we are passionate about ensuring all our customers around the globe have access to all the content they need, when they need it. IMDb sits at the intersection of the entertainment, media, and technology markets inside the world’s most innovative and consumer-centric company – Amazon.com. IMDb employees enjoy the benefits of working for Amazon with the autonomy of working on a smaller, nimble team.&lt;br /&gt;
&lt;br /&gt;
As an Ontologist, you work as part of a global team to deliver world-class, intuitive, and comprehensive taxonomy and ontology models to optimize product delivery for IMDb web and mobile experiences. You collaborate with business partners and engineering teams to deliver knowledge-based solutions to enable discovery and engagement with IMDb content. In this role you will directly impact the customer experience as well as the company&#039;s product knowledge foundation across all customer cohorts.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Specific responsibilities include the following:&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Develop logical, semantically rich, and extensible data models for IMDb&#039;s expansive entertainment catalog.&lt;br /&gt;
&lt;br /&gt;
Ensure our ontologies provide comprehensive sub-domain coverage that are available for machine ingestion and inference&lt;br /&gt;
&lt;br /&gt;
Research worldwide understanding of entertainment content to develop scalable data models that solve customer problems and enhance entity discovery&lt;br /&gt;
&lt;br /&gt;
Contribute to the development of new tools, features and processes for the Ontology team&lt;br /&gt;
&lt;br /&gt;
Support the content expansion team focused on driving the overall IMDb Content product and business strategy and execution.&lt;br /&gt;
&lt;br /&gt;
BASIC QUALIFICATIONS&lt;br /&gt;
*Experience working in ontology and/or taxonomy roles&lt;br /&gt;
*Proven skills in data retrieval and data research techniques&lt;br /&gt;
*Ability to quickly understand complex processes and communicate them in simple language&lt;br /&gt;
*Ability to communicate knowledge-based requirements and needs to engineering and retail teams&lt;br /&gt;
*Familiarity with Semantic Web technologies (RDF/s, OWL), query languages (SPARQL) and validation/reasoning standards (SHACL, SPIN)&lt;br /&gt;
*Detail-oriented problem-solving, ability to work in fast-changing environment and manage ambiguity&lt;br /&gt;
*Proven track record of strong communication and interpersonal skills&lt;br /&gt;
*Proficient English language skills&lt;br /&gt;
&lt;br /&gt;
PREFERRED QUALIFICATIONS&lt;br /&gt;
Master’s degree in Library Science, Information Systems, Linguistics or other relevant fields&lt;br /&gt;
&lt;br /&gt;
Experience building ontologies in the entertainment and semantic search spaces&lt;br /&gt;
&lt;br /&gt;
Experience working with schema-level constructs (e.g. higher-order classes, punning, property inheritance)&lt;br /&gt;
&lt;br /&gt;
Proficiency in SQL, SPARQL&lt;br /&gt;
&lt;br /&gt;
Familiarity with software engineering life cycle&lt;br /&gt;
&lt;br /&gt;
Familiarity with ontology manipulation programming libraries&lt;br /&gt;
&lt;br /&gt;
Exposure to data science and/or machine learning, including graph embeddings&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Amazon is an equal opportunities employer. We believe passionately that employing a diverse workforce is central to our success. We make recruiting decisions based on your experience and skills. We value your passion to discover, invent, simplify and build. Protecting your privacy and the security of your data is a longstanding top priority for Amazon. Please consult our Privacy Notice (https://www.amazon.jobs/en/privacy_page) to know more about how we collect, use and transfer the personal data of our candidates.&lt;br /&gt;
&lt;br /&gt;
==KNOWLEDGE GRAPH AND SEMANTICS SME==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;About the job&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Company: Solutions Driven &lt;br /&gt;
&lt;br /&gt;
Location: United States &lt;br /&gt;
&lt;br /&gt;
Type: Remote &lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;About the team:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Our client&#039;s Data Team helps to transform every aspect of their business. We are highly skilled at formulating data strategy, defining business and technology initiatives across the data management lifecycle, and aligning multi-year strategic roadmaps with client’s business goals. As digital technologies advance and regulations tighten, today’s consumers – and, therefore, today’s businesses – are becoming more aware of the importance of good quality data. We work to establish holistic ways to effectively manage data through the modern data supply chain and facilitate consumption through analytics, modelling, AI, machine learning, dashboarding, and reporting.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;What You’ll Get to Do:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The practice leader will involve shaping and executing the hiring strategy, developing the best of in-house talent, leading talent, and collaboration with the others in the data &amp;amp; analytics practice to bring cutting edge solutions to the market.&lt;br /&gt;
&lt;br /&gt;
Graph and Semantic Engineering are emerging enterprise data technologies that leverage ontologies and related advanced analytics – including artificial intelligence – to build and represent knowledge in expressive and novel ways&lt;br /&gt;
&lt;br /&gt;
Primary responsibility in this role is to help clients succeed by understanding, exploiting, designing, and implementing solutions involving graph and semantic technologies –  Solutions Driven United States Remote including advanced data management and analytics, and other innovative new ways&lt;br /&gt;
&lt;br /&gt;
In the fast-changing landscape of potential partners in graph technology, an ongoing curiosity and developing connections with key partners, will enable valuable support for clients&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;What You’ll Bring with You:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Graph and Semantic Data Consultants will require multi-disciplinary capabilities required to explain, design and create powerful knowledge graph models – a good balance of business and technical understanding&lt;br /&gt;
&lt;br /&gt;
Sound understanding of Knowledge Representation and Semantic Technologies (OWL, RDF, SWRL, SPARQL, JSON-LD) including semantic modelling and data integration, data unification, knowledge graph design, ontology and taxonomy&lt;br /&gt;
&lt;br /&gt;
Awareness of the breadth of use cases for these technologies, prudent user interface design which exploit the power of graph and mitigates the challenges of visualizing graph data&lt;br /&gt;
&lt;br /&gt;
Experience with Graph/Triple Stores and related technologies, e.g. PoolParty, Stardog, TigerGraph, Ontotext GraphDB, Grakn, MarkLogic, Metaphactory, Neo4j&lt;br /&gt;
&lt;br /&gt;
Understanding of data/information architecture and modeling languages such as UML, ER, IDEF, Data Flow Diagram&lt;br /&gt;
&lt;br /&gt;
Excellent Communication skills – be an effective, passionate, trusted advocate and communicator for Graph and Semantic Technologies&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;Other desired skills:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Good awareness of financial services domain is a great advantage&lt;br /&gt;
* Ability to utilize NLP, machine learning and semantic text mining to translate data into machine understandable representations to support advanced analytics, classification, inference processing and knowledge extraction&lt;br /&gt;
* Familiarity with mapping techniques, e.g. R2RML, for transforming relational/tabular datasets into triples&lt;br /&gt;
* Agile software development frameworks, e.g. Scrum, Kanban, Extreme Programming&lt;br /&gt;
* Understanding of data management and governance toolsets and methodologies&lt;br /&gt;
* Experience in applying semantic technologies in an industry setting, or practical experience as part of an academic, Library Science or Information/Knowledge &lt;br /&gt;
* Management career &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;Why Us?&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We are the largest Financial Services focused consultancy in the world, serving everyone from global banks to emerging FinTechs, from strategy through digital transformation, design, business consulting, data and analytics, cyber, cloud, technology architecture, and engineering. We are young and growing firm. We maintain an entrepreneurial spirit and growth mindset, and have minimal bureaucracy. We have no internal silos that get in the way of your career opportunities or ability to focus on our clients and make a difference to the business. We offer the opportunity for everyone to learn rapidly, take on tough challenges, and get promoted quickly. We take pride in our creative, collaborative, diverse, and inclusive culture, where everyone can Be Yourself at Work. We offer highly competitive benefits, including medical, dental and vision insurance, a 401(k) plan, tuition reimbursement, and a work culture focused on innovation and creation of lasting value for our clients and employees.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;Ready to take the Next Step&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If this sounds like you, we would love to hear from you. This is an opportunity to make a difference and contribute to a highly successful company with a significant growth trajectory.&lt;br /&gt;
&lt;br /&gt;
This position may be performed remotely anywhere within the United States except the State of Colorado.&lt;br /&gt;
&lt;br /&gt;
Apply here: https://www.linkedin.com/jobs/view/2943587552/&lt;br /&gt;
&lt;br /&gt;
==Semantics Specialist==&lt;br /&gt;
&lt;br /&gt;
Company: [A]&lt;br /&gt;
&lt;br /&gt;
Location: 100% Remote from North America (&#039;&#039;&#039;excluded:&#039;&#039;&#039; CA, MA or NY)&lt;br /&gt;
&lt;br /&gt;
[A] is looking for a Semantics Specialist to help make the world smarter with intelligent content as a part of the [A] team. [A] seeks motivated, creative, innovative candidates eager to start a client-facing role with a mix of taxonomy, thesaurus, and ontology development experience with an interest in learning or growing their experience in content development and content modeling. &lt;br /&gt;
&lt;br /&gt;
The ideal candidate will have experience and knowledge working on the following content principles:&lt;br /&gt;
&lt;br /&gt;
Content Semantics (controlled vocabularies, ontology, taxonomy, metadata, schemas etc.). Content Structure (content modeling, markup, schema, DITA, XML) This role is NOT focused on Content &amp;quot;Interchange&amp;quot; (metadata, XSLT, XML Processors and Parsers), or Content Operations (CMS tools and platforms management, Content acquisition tools, systems, processes, and patterns). At [A], these practices have their own speciality focus. Having familiarity with these related practices is very helpful. &lt;br /&gt;
&lt;br /&gt;
The responsibilities for this role deeply involve the principles of content semantics and structure. And this role requires one to build an understanding of the application of semantics principals at abstract, architecture, and at very tactical applied levels within enterprise knowledge and customer experience delivery environments.&lt;br /&gt;
&lt;br /&gt;
This is a remote job that can be done mostly from your home office anywhere in North America (we are unable to consider candidates in the states of CA, MA or NY at this time), but may involve occasional travel to a client location (once safe to do so). &lt;br /&gt;
&lt;br /&gt;
The Semantics Specialist will work within the larger Content Intelligence practice to build next-generation content systems for clients, helping them craft intelligent content for multi-channel marketing and increase the value of semantically-rich content. You will participate in, or lead, client projects that include development of semantic models, taxonomies, thesauri, controlled vocabularies, and formal ontologies in conjunction with content structures and content tools experts.&lt;br /&gt;
&lt;br /&gt;
You’ll define content relationships by using taxonomies, ontologies, and metadata. This includes implementation of semantic models within semantic software tools such as those from Synaptica, the Semantic Web Company, and other semantic provider solutions that are fitting into new Content Intelligence architectures.&lt;br /&gt;
&lt;br /&gt;
Link: http://jobs.simplea.com/apply/cTn4Ye5FlK/Semantics-Specialist&lt;br /&gt;
&lt;br /&gt;
==Head of Ontologies and Business Domain Data Modeling - NJ USA==&lt;br /&gt;
&lt;br /&gt;
Date: 09/2020&lt;br /&gt;
&lt;br /&gt;
Location: East Hanover, New Jersey, USA&lt;br /&gt;
&lt;br /&gt;
Strategic purpose of this role is to organize Novartis data, make it easily accessible, and useful to authorized roles. This role will help drive the development and execution of Novatis’ ambition to turn data into a real strategic asset across the organization. This ambition is one of key pillars in the broader digital transformation happening at Novartis to be a ‘medicines and data science company.’ This data centric role will facilitate using data to digitize the biopharma value chain in finding right target, right tissue, right safety, right patients, right commercial potential and to optimize omni channel stakeholder experience and engagement. More specifically, the purpose of this role is to engineer, in partnership with over 10+ business units and Technology partners (internal and external), global adoption of consistent Data Life Cycle (DLC) management framework, processes, architecture and supporting tools. Key areas in the Data Life Cycle management includes Master Data Management, Data Governance, Information Modeling, and Data Science Enablement. This role will specifically focus on improving maturity of Data Ownership and Capability Building in all bio pharma data domains across the value chain; omics, compounds, diseases, patients, payers, providers, sites, trials, KOLs, HCPs, EMR, EHR, patient journeys, employees, contracts, vendors, products, etc&lt;br /&gt;
&lt;br /&gt;
This role will directly report and work with the Head of Data Strategy in the Group Digital Office.&lt;br /&gt;
&lt;br /&gt;
URL: https://careers.iscb.org/jobs/view/7251&lt;br /&gt;
&lt;br /&gt;
==Head of Knowledge Graph &amp;amp; Semantics - NJ USA==&lt;br /&gt;
&lt;br /&gt;
Date: 09/2020&lt;br /&gt;
&lt;br /&gt;
Location: East Hanover, New Jersey, USA&lt;br /&gt;
&lt;br /&gt;
Strategic purpose of this role is to organize Novartis data, make it easily accessible, and useful to authorized roles. This role will help drive the development and execution of Novartis’ ambition to turn data into a real strategic asset across the organization. This ambition is one of key pillars in the broader digital transformation happening at Novartis to be a ‘medicines and data science company.’ This data centric role will facilitate using data to digitize the biopharma value chain in finding right target, right tissue, right safety, right patients, and right commercial potential and to optimize Omni channel stakeholder experience and engagement.&lt;br /&gt;
&lt;br /&gt;
More specifically, the purpose of this role is to drive, in partnership Digital Data Science and AI teams, over 10+ business units, Technology partners (internal and external) adoption of consistent framework, processes, architecture and supporting tools for analytics needs. This role will help create high quality data assets to enable analytics using data across all bio pharma data domains and value chain; omics, compounds, diseases, patients, payers, providers, sites, trials, KOLs, HCPs, EMR, EHR, patient journeys, employees, contracts, vendors, products, etc.&lt;br /&gt;
&lt;br /&gt;
This role will directly report and work with the Head of Data Strategy in the Group Digital Office.&lt;br /&gt;
&lt;br /&gt;
URL: https://careers.iscb.org/jobs/view/7250&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Software Engineer to work on the Data Team at Financial Services Firm NYC/NJ==&lt;br /&gt;
Date: 08/2020&lt;br /&gt;
&lt;br /&gt;
Location: NYC /NJ&lt;br /&gt;
&lt;br /&gt;
Software Engineer to work on the data team for a premier financial services firm in NYC/NJ on Enterprise Data/Historical Data/Metadata / Semantic Wikis/SPARQL/OWL – Ontology Semantic Web Language/RDFS.&lt;br /&gt;
&lt;br /&gt;
status: filled&lt;br /&gt;
&lt;br /&gt;
==Analytical Linguist, Knowledge Engine Google - San Francisco, CA, USA ==&lt;br /&gt;
Behind Google&#039;s Knowledge Engine (KE) team is the world&#039;s largest and most comprehensive semantic graph. We power features across Google, including Ads, YouTube, Google Play, Geo, and many others. Growing Knowledge Graph requires precise models describing how entities fit together. As an Analytical Linguist on the Knowledge Engine team, you will analyze graph structures and content, develop new semantic representations, and work with providers and consumers of data to guide the development and usage of knowledge structures. Using a variety of semantic modeling techniques, you’ll work to improve entity accessibility. You will be responsible for judging tradeoffs between formality and usability and deciding when to borrow from an existing ontology and when to develop new structures.&lt;br /&gt;
&lt;br /&gt;
Responsibilities&lt;br /&gt;
Analyze graph structures and content.&lt;br /&gt;
Develop new semantic representations.&lt;br /&gt;
Work with providers and consumers of data to guide the development and usage of knowledge structures to improve entity accessibility.&lt;br /&gt;
Use a variety of sematic modeling techniques, judging tradeoffs between formality and usability and deciding when to borrow from an existing ontology and when to develop new structures.&lt;br /&gt;
&lt;br /&gt;
Minimum qualifications&lt;br /&gt;
MA/MS in linguistics, computer science, library/information science, philosophy, or a related field, or equivalent work experience.&lt;br /&gt;
Coding experience in Python, C/C++, Java, or Go.&lt;br /&gt;
Preferred qualifications&lt;br /&gt;
PhD in relevant field, or equivalent work experience.&lt;br /&gt;
Experience with ontology development (RDF/OWL, SPARQL, the Semantic Web or Frame based KR systems).&lt;br /&gt;
Area&lt;br /&gt;
There is always more information out there, and the Knowledge team has a never-ending quest to find it and make it accessible. We&#039;re constantly refining our signature search engine to provide better results, and developing offerings like Google Instant, Google Voice Search and Google Image Search to make it faster and more engaging. We&#039;re providing users around the world with great search results every day, but at Google, great just isn&#039;t good enough. We&#039;re just getting started.&lt;br /&gt;
&lt;br /&gt;
Team or role:Program Management&lt;br /&gt;
Job type:Full-time&lt;br /&gt;
Last updated: Dec 31, 2014&lt;br /&gt;
Job location(s):San Francisco, CA, USA&lt;br /&gt;
&lt;br /&gt;
https://www.google.com/about/careers/search#!t=jo&amp;amp;jid=86625001&amp;amp;&lt;br /&gt;
&lt;br /&gt;
==Wiss. Mitarbeiterin / Wiss. Mitarbeiter befristet auf 6 Monate (E 13 TV-L FU)==&lt;br /&gt;
&lt;br /&gt;
Im Rahmen der Juniorprofessur für Web Science/Human-Centered Computing (Prof. Dr. Claudia Müller-Birn) an der FU Berlin wurde im letzten Jahr ein Team aufgebaut, das schwerpunktmäßig in den Bereichen der Web-basierten Wissensgenerierung und des Digitalen Lernens tätig ist. Die Gruppe arbeitet mit Universitäten und Forschungseinrichtungen sowie non-profit Organisationen national und international zusammen.&lt;br /&gt;
&lt;br /&gt;
Weitere Informationen unter: http://www.mi.fu-berlin.de/inf/groups/hcc/index.html&lt;br /&gt;
&lt;br /&gt;
Das Aufgabengebiet umfasst die Mitarbeit in dem Forschungsprojekt Neonion – Collaborative Annotation in Humanities and Arts –, innerhalb dessen eine semantische Annotationssoftware in einem nutzerzentrierten Designprozesses weiterentwickelt wird. In einem Pilotprojekt mit einem Projektpartner aus dem Bereich der Geisteswissenschaften werden neuartige Interaktionskonzepte zur semantischen Annotation von Dokumenten erforscht. Eine aktive Beteiligung an laufenden Forschungsanträgen und Publikationsaktivitäten wird erwartet. Eine Möglichkeit zur Verlängerung dieser zeitlich befristeten Tätigkeit ist je nach Verfügbarkeit einer Folgefinanzierung gegeben.&lt;br /&gt;
&lt;br /&gt;
Einstellungsvoraussetzung ist ein abgeschlossenes Hochschulstudium in Informatik oder Wirtschaftsinformatik (Diplom oder Master). Alternativ ist ein Abschluss in einem verwandten Bereich mit entsprechender Erfahrung in der Informatik möglich.&lt;br /&gt;
&lt;br /&gt;
Folgende Kenntnisse sind für die Besetzung der Stelle von Vorteil und wünschenswert:&lt;br /&gt;
• Kenntnisse in den Bereichen der Semantischen Technologien und Web Technologien&lt;br /&gt;
• Kenntnisse im Bereich der nutzerzentrierten Softwareentwicklung und Usability&lt;br /&gt;
• Erfahrung im wissenschaftlichen Arbeiten und insbesondere in qualitativen Forschungsmethoden&lt;br /&gt;
• Kommunikationsfähigkeit und eigenständiges Arbeiten&lt;br /&gt;
• Erfahrungen im Umgang mit Confluence and Jira&lt;br /&gt;
• Bereitschaft zur interdisziplinären Zusammenarbeit in einem internationalen Konsortium&lt;br /&gt;
• Sehr gute Englische Kenntnisse in Wort und Schrift&lt;br /&gt;
&lt;br /&gt;
Bitte senden Sie Ihre Bewerbung per E-Mail bis zum 30.06.2014 als eine einzige PDF-Datei mit einem ausführlichen Lebenslauf (inkl. Kenntnisse und praktischen Erfahrungen), Kopien Ihrer akademischen Abschlüsse, Referenzschreiben (falls vorhanden) und ein Motivationsschreiben an Prof. Dr. Cl. Müller-Birn (clmb@inf.fu-berlin.de).&lt;br /&gt;
&lt;br /&gt;
Für weitere Informationen zu dieser Ausschreibung wenden Sie sich bitte ebenfalls per E-Mail an Prof. Dr. Cl. Müller-Birn (clmb@inf.fu-berlin.de).&lt;br /&gt;
&lt;br /&gt;
==Wissenschaftliche/r Mitarbeiter/in (Data Scientist) am Museum für Naturkunde Berlin (MfN)==&lt;br /&gt;
&lt;br /&gt;
http://www.naturkundemuseum-berlin.de/fileadmin/startseite/institution/stellenausschreibungen/17_2014.pdf&lt;br /&gt;
&lt;br /&gt;
Arbeitszeit: 100 v.H. d. regelm. wöchentlichen Arbeitszeit&lt;br /&gt;
&lt;br /&gt;
Befristung: Drittmittelfinanzierung, zum nächstmöglichen Zeitpunkt befristet bis zum 31. Mai 2015 &lt;br /&gt;
&lt;br /&gt;
Entgeltgruppe: E13 TV-L Berlin &lt;br /&gt;
&lt;br /&gt;
Aufgabengebiet: Die erfolgreiche Bewerberin / der erfolgreiche Bewerber soll das Museum für Naturkunde in zwei Projekten des Forschungsbereichs „Digitale Welt und Informationsmanagement“ unterstützen: Das EU-geförderte Projekt EuropeanaCreative (www.pro.europeana.eu/web/europeana-creative) befasst sich mit den Möglichkeiten der Nachnutzung digitaler Medien über die Plattform Europeana (www.europeana.eu). Das DFG-geförderte Projekt „German Federation for the Curation of Biological Data” (www.gfbio.org) arbeitet am Aufbau einer grundlegenden Forschungsinfrastruktur, die den Austausch von biologischen und umweltbezogenen Forschungsdaten ermöglichen bzw. vereinfachen soll. Nach der derzeitigen 18-monatigen Pilotphase sind zwei je 36 Monate laufende Anschlussprojekte geplant. &lt;br /&gt;
&lt;br /&gt;
Die Stelle eignet sich als Startpunkt für den Aufbau einer Forschergruppe am MfN mit Fokus auf semantischer Biodiversitätsinformatik. Das Ziel des Forschungsbereichs „Digitale Welt und Informationsmanagement“ ist die Implementierung einer zukunftsfähigen Architektur für naturkundliche Daten und Datenrepositorien, welche neue Möglichkeiten zur Beantwortung von Forschungsfragen sowohl im Bereich „Data Science“ als auch „Public Engagement with Science“ eröffnet.&lt;br /&gt;
&lt;br /&gt;
Der Aufgabenbereich der erfolgreichen Bewerberin / des erfolgreichen Bewerbers umfasst:&lt;br /&gt;
* Methoden und Prozesse von Linked Open Data, semantischem Web und Wissensmanagement in Hinblick auf ihre Nutzbarkeit für naturkundliche * Daten zu evaluieren und exemplarisch zu testen bzw. umzusetzen&lt;br /&gt;
* Grundlagen für eine effektivere Zusammenarbeit von Forschern verschiedener Fachrichtungen, z.B. Taxonomie, Umweltwissenschaften und  Naturschutz, mittels semantischen Wissensmanagements zu erarbeiten&lt;br /&gt;
* Möglichkeiten von „Data-driven-science“ und „Open Science“ in Pilotprojekten zu demonstrieren&lt;br /&gt;
* konzeptionelle und strategische Weiterentwicklung von GFBio in Zusammenarbeit mit dem GFBio Team am MfN voranzutreiben&lt;br /&gt;
* in Kooperation mit GFBio Projektpartnern am Aufbau des GFBio Terminologieservers mitzuwirken und am MfN vorhandene Terminologien bzw. Ontologien für die Verwendung in GFBio zu erschließen&lt;br /&gt;
* Erschließung, Analyse und Impact naturkundlicher Medien von Europeana Creative mit Mitteln des Semantic Webs zu verbessern&lt;br /&gt;
* Unterstützung der Vermittlung der Projektergebnisse nach innen und außen durch Kommunikation, Präsentationen, Berichte und Publikationen&lt;br /&gt;
&lt;br /&gt;
Anforderungen:&lt;br /&gt;
* abgeschlossenes wissenschaftliches Hochschulstudium der Informatik, Bioinformatik oder verwandter Fächer&lt;br /&gt;
* Nachweis aktiver wissenschaftlicher Publikationstätigkeit&lt;br /&gt;
* praktische Erfahrungen im Bereich Linked Open Data, Semantic Web oder Ontologieentwicklung&lt;br /&gt;
* sehr gute Kenntnisse der deutschen und englischen Sprache in Wort und Schrift&lt;br /&gt;
* ausgezeichnete Kommunikations- und Teamfähigkeit, insbesondere in interdisziplinären Teams&lt;br /&gt;
* selbstständiges und strukturiertes Arbeiten&lt;br /&gt;
* Freude an der Erprobung innovativer Vermittlungswege für wissenschaftliche bzw. naturkundliche Fragestellungen&lt;br /&gt;
* überdurchschnittliche Einsatzbereitschaft und Zuverlässigkeit&lt;br /&gt;
&lt;br /&gt;
Von Vorteil sind:&lt;br /&gt;
* Promotion&lt;br /&gt;
*Erfahrung im Informations- und Wissensmanagement an wissenschaftlichen Einrichtungen&lt;br /&gt;
*Erfahrung in der Anleitung von Teams&lt;br /&gt;
*Erfahrungen mit Semantic Media Wiki oder Wikidata&lt;br /&gt;
*Kenntnisse der Aufgaben und Arbeitsweisen naturkundlicher Forschungsmuseen&lt;br /&gt;
&lt;br /&gt;
==Data Analytics Technical Architect - Chase Auto Finance - Jersey City, NJ==&lt;br /&gt;
Chase is a leader in the financial services industry, providing banking, mortgages, credit cards, loans, payment processing and investment services to 50 million customers - 1 out of every 6 Americans. As a division of JPMorgan Chase &amp;amp; Co. (NYSE:JPM), we:&lt;br /&gt;
Serve 21 million households with consumer banking relationships&lt;br /&gt;
Lent $17 billion to small businesses in 2011&lt;br /&gt;
Are one of the nation&#039;s largest credit card issuers, with more than 64 million credit cards in circulation&lt;br /&gt;
Service 8 million mortgage and home equity loans&lt;br /&gt;
While we operate across a broad range of businesses, our mission at Chase is quite simple: to be the industry leader in customer service. Our employees put the firm&#039;s resources to work every day for our customers.&lt;br /&gt;
 &lt;br /&gt;
Chase offers a dynamic environment and the training and support to meet your full potential. Our company is widely recognized as a great place to work, to grow and to invest for the future. Join our team.&lt;br /&gt;
The Chase Auto Finance (CAF) Data Management Program is a multi-year technology initiative to support CAF and enterprise data management processes addressing operational, financial, regulatory and analytical reporting requirements.  The program includes the long term development of a strategic platform for CAF with data sourcing, data enrichment, analytical, monitoring and reporting capabilities.  This platform will also integrate with cross-LOB efforts to create a single consolidated customer view.&lt;br /&gt;
 &lt;br /&gt;
The CAF Data Management team is looking for a highly motivated individual that will have a strong foundational knowledge and experience with distributed systems and computing systems with hands-on engineering skills. Experience with a range of big data architectures and broad understanding and experience of real-time analytics, NoSQL data stores, big data analytics products. In addition, experience with data modeling and data management, analytical tools, languages, or libraries is required.&lt;br /&gt;
The role is heavily collaborative, working closely with several internal business units as well as a firm-wide organization driving standards and best practices and the candidate will have several years experience working in client-focused roles.&lt;br /&gt;
 &lt;br /&gt;
Key Areas of Responsibility:&lt;br /&gt;
This individual will be responsible for guiding the full lifecycle of a Big Data Analytics solution, including requirements analysis, platform selection, technical architecture design, application design and development, testing, and deployment.  We are looking for candidates with a broad set of technology skills to be able to design and build robust solutions for big data problems and learn quickly as the platform grows.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Hands-on experience with the Hadoop stack (e.g. MapReduce, Sqoop, Pig, Hive, Hbase, Flume)&lt;br /&gt;
Hands-on experience with related/complementary open source software platforms and languages (e.g. Java, Linux, Apache, Perl/Python/PHP, Chef)&lt;br /&gt;
Hands-on experience with ETL (Extract-Transform-Load) tools (e.g Informatica,  Talend, Pentaho)&lt;br /&gt;
Hands-on experience with analytical tools, languages, or libraries (e.g. SAS, SPSS, R, Mahout)&lt;br /&gt;
Hands-on experience with &amp;quot;productionalizing&amp;quot; Hadoop applications (e.g. administration, configuration management, monitoring, debugging, and performance tuning)&lt;br /&gt;
Previous experience with high-scale or distributed RDBMS (Teradata, Netezza, Greenplum, Aster Data, Vertica)&lt;br /&gt;
Knowledge of NoSQL platforms (e.g. key-value stores, graph databases, RDF triple stores)&lt;br /&gt;
Experience with at least 4 of the following activities in the context of high-scale or distributed systems:&amp;lt;br&amp;gt;&lt;br /&gt;
1.    Implementation of ETL applications&amp;lt;br&amp;gt;&lt;br /&gt;
2.    Implementation of reporting applications&amp;lt;br&amp;gt;&lt;br /&gt;
3.    Application/implementation of custom analytics&amp;lt;br&amp;gt;&lt;br /&gt;
5.    Data migration from existing data stores&amp;lt;br&amp;gt;&lt;br /&gt;
6.    Infrastructure and storage design&amp;lt;br&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
Bachelors Degree&lt;br /&gt;
Minimum 6 years systems development and implementation experience&lt;br /&gt;
Minimum 5+ years of hands on Java experience building scalable solutions&lt;br /&gt;
Minimum 2 years of hands-on experience with programming on high-scale or distributed systems (such as Hadoop)&lt;br /&gt;
Minimum 3+ years experience with full application development lifecycle (analysis, architecture, design, development, testing and deployment).&lt;br /&gt;
Minimum 3+ years experience working with multiple RDBMS and associated applications&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Preferred Skills&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Knowledge of the Auto Loan Industry&lt;br /&gt;
Experience in designing and implementation of Enterprise Information Integration architectures&lt;br /&gt;
An understanding of Business Intelligence concepts including Dashboards, and the data requirements of senior executives&lt;br /&gt;
Knowledge of Teradata Aster&lt;br /&gt;
Knowledge of Cognos&lt;br /&gt;
&lt;br /&gt;
JPMorgan Chase is an Equal Opportunity and Affirmative Action Employer, M/F/D/V&lt;br /&gt;
&lt;br /&gt;
https://jpmchase.taleo.net/careersection/2/jobdetail.ftl?lang=en&amp;amp;job=1239446&amp;amp;src=JB-13027&lt;br /&gt;
&lt;br /&gt;
==Senior Taxonomist in Financial Services New York, New York==&lt;br /&gt;
&lt;br /&gt;
Role Description:&lt;br /&gt;
&lt;br /&gt;
This role requires someone with a strong background in system and business analysis.  Key requirement of this position is strong communication skills both written and verbal, as the successful candidate will be expected to deal directly with business stakeholders to elicit requirements and clearly present these to the technical team for development.  Demonstration of requirements gathering, modelling and documentation would be required. &lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Responsibilities Include:&lt;br /&gt;
&lt;br /&gt;
- Developing, mapping, evaluating, and/or maintaining taxonomies&lt;br /&gt;
&lt;br /&gt;
- Conducting content audits and content analysis&lt;br /&gt;
&lt;br /&gt;
- Developing metadata schemas, tagging, conducting metadata audits&lt;br /&gt;
&lt;br /&gt;
- Conducting user research, usability testing Interviewing stakeholders, subject matter experts, and end users&lt;br /&gt;
&lt;br /&gt;
- Providing training on taxonomy maintenance, tagging, or tool use&lt;br /&gt;
&lt;br /&gt;
- Facilitating working sessions&lt;br /&gt;
&lt;br /&gt;
- Support taxonomy implementation&lt;br /&gt;
&lt;br /&gt;
- Support integration of content-, document-, taxonomy-management tools, search engines etc.&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Required Skills and Qualifications&lt;br /&gt;
&lt;br /&gt;
- Experience in working in document management space&lt;br /&gt;
&lt;br /&gt;
- Experience with--or knowledge of--taxonomy, metadata, controlled vocabularies, and classification&lt;br /&gt;
&lt;br /&gt;
- Excellent attention to detail&lt;br /&gt;
&lt;br /&gt;
-Very strong communication, both written and verbal.&lt;br /&gt;
&lt;br /&gt;
-Ability to lead/direct small focused groups of individuals.&lt;br /&gt;
&lt;br /&gt;
-Experience in dealing with key business stakeholders in a professional manner including client engagement and building relationships with business and partnering technology teams.&lt;br /&gt;
&lt;br /&gt;
contact: info@kona.llc&lt;br /&gt;
&lt;br /&gt;
==Software Engineer In The Semantic Algorithms And News Delivery==&lt;br /&gt;
Location:  Roseland, NJ&lt;br /&gt;
&lt;br /&gt;
Date: 29 September 2012&lt;br /&gt;
&lt;br /&gt;
Acquire Media analyzes and distributes the news. We have an opening for a software engineer to help our semantic algorithms and news delivery teams.&lt;br /&gt;
&lt;br /&gt;
Here&#039;s the link: (on Craigslist) http://newjersey.craigslist.org/sof/3409136469.html&lt;br /&gt;
&lt;br /&gt;
== Senior Java Developer With Extensive Semantic Web Experience - Nature Publishing Group ==&lt;br /&gt;
&lt;br /&gt;
Location: London, UK&amp;lt;br&amp;gt;&lt;br /&gt;
Date: 27 November 2012&lt;br /&gt;
&lt;br /&gt;
Nature Publishing Group - the prestigious international scientific publishing company - seeks an enthusiastic Software Developer to work&lt;br /&gt;
in our London office as an integral part of the team developing our growing portfolio of online products.&lt;br /&gt;
&lt;br /&gt;
You will enhance and expand products across http://www.nature.com, ensuring that your work is well designed, maintainable and tested. You will fit this description: http://bit.ly/hacker-def and enjoy having conversations about new frameworks, new technologies and the joys of distributed version control systems.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The successful candidate must:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
have great technical ability in the field of web development write high quality code delight in having an intimate understanding of the internal workings of a system be comfortable talking about design, code and the requirements that shape it enjoy the intellectual challenge of creatively overcoming or circumventing limitations have a great desire to learn, share knowledge and crave constructive criticism&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;You will also be able to demonstrate the required skills:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
extensive experience of Java 6, dependency injections frameworks and relational databases experience in web servers (Jetty, Tomcat), XML, web services, JUnit, Mockito and JPA knowledge of triplestore and/or graphstores, semantic technologies, RDF and SPARQL knowledge of XML Database (MarkLogic), XQuery experience of bug tracking, distributed version control systems and continuous integration a good understanding of software development methodologies&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Nature Publishing Group (NPG) is a division of Macmillan Publishers Ltd, dedicated to serving the academic, professional scientific and&lt;br /&gt;
medical communities. This role is in a great location in King&#039;s Cross with a lovely working environment, option for macbook, free canteen,&lt;br /&gt;
exciting and challenging work, enthusiastic colleagues, and a sensible work/life balance. The salary is excellent and negotiable for the&lt;br /&gt;
right person.&lt;br /&gt;
&lt;br /&gt;
contact marco.neumann@gmail.com&lt;br /&gt;
&lt;br /&gt;
==Sales Executives at Neo Technology, Germany ==&lt;br /&gt;
&lt;br /&gt;
Location&lt;br /&gt;
Germany (Munich)&lt;br /&gt;
&lt;br /&gt;
Overview&lt;br /&gt;
As a Sales Executive you are a hunter and closer with experience selling to developers, architects, IT Directors, and consultants within F100 firms, government agencies, and to start-ups depending on your experience. Creative, energetic and a self-starter, you understand the sales process, how to sell innovation and disruption through customer vision expansion. You can drive deals forward and compress decision cycles. You also love understanding a product in depth and then communicating that product to the marketplace. If you’re a high achiever, we want to talk to you.&lt;br /&gt;
&lt;br /&gt;
We’re looking for a track record of proven sales success with 3 -15 years of experience in software sales.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements, Qualifications, Skills &amp;amp; Abilities&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
*Qualify inbound leads representing key stakeholders, CIOs, CTO’s, IT Director’s, IT architect’s and consultant’s from various organizations&lt;br /&gt;
*Demonstrating Neo4j’s capabilities over the Web and phone&lt;br /&gt;
*Meeting/exceeding activity, pipeline, and booking targets&lt;br /&gt;
*Discipline in documenting various details including use case, purchase timeframes, next steps, and forecasting&lt;br /&gt;
*Ensuring 100% satisfaction with all customers and partners&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
*B.Sc. or M.Sc. degree required&lt;br /&gt;
*Obvious passion and people skills&lt;br /&gt;
*Extensive prior customer relationships&lt;br /&gt;
*A history of consistent over-achievement&lt;br /&gt;
*Experience in opportunity creation and the challenges of building a business&lt;br /&gt;
&lt;br /&gt;
[http://www.neotechnology.com/2012/10/neo-technology-is-hiring-sales-executives-germany/ Apply Here]&lt;br /&gt;
&lt;br /&gt;
==Software Developer with Semantic Web experience (University of Colorado, Boulder, CO)==&lt;br /&gt;
The Faculty Information System team at the University of Colorado Boulder is searching for a software developer who loves data and learning. Someone who would enjoy working on the beautiful Boulder campus, minutes from outdoor activities and Downtown Boulder. If this describes you, or someone you know, please read on.&lt;br /&gt;
&lt;br /&gt;
Our team is launching VIVO CU-Boulder, a key component of CU-Boulder&#039;s presence on the next generation Web that harnesses the possibilities of Linked Data and Open Data. The FIS team is a small, collaborative group employing agile and lean software development practices to deliver service-oriented solutions in an expanding environment. We are looking for someone who enjoys working closely with others, and who would be committed to the continuous improvement of FIS and our software development process. Candidates for the position should demonstrate success in data management and web application development using current methods and technologies. Experience with VIVO, W3C Semantic Web technologies, and/or prior contributions to an open source community are a plus.&lt;br /&gt;
&lt;br /&gt;
This position will be open until filled. Applications submitted by Wednesday June 13, 2012 will be given full consideration. For more information on this position and the full benefits package, please see http://www.jobsatcu.com/applicants/Central?quickFind=68914&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Client Development Manager (Reuters Media)==&lt;br /&gt;
 &lt;br /&gt;
Our engineering team develops business-critical applications and services for the largest news agency in the world.  We provide publishers and broadcasters with news text, photos, and video over both  Internet and satellite.  As part of a small team of high-performers, you will be our principal client-facing technical expert, and will help guide development for our products.  The ideal candidate has a deep understanding of publisher workflows, is a talented programmer, and has experience with pre- and post-sales support. &lt;br /&gt;
 &lt;br /&gt;
As our client-facing technical expert, you understand and can evangelize technical standards such as NewsML-G2 and RightsML, and keep abreast of the latest trends in publishing.  You can work with our metadata experts and provide them with our client’s perspective.  You also have an opportunity to interact with IPTC leadership in the development of new standards such as rNews.  Participation in New York City industry groups such as the NY Tech Meetup, Hacks and Hackers, and the New York Semantic Web group are all a big plus.&lt;br /&gt;
 &lt;br /&gt;
Responsibilities:&lt;br /&gt;
* Advise and assist clients on making the best use of our platform and data, including occasional visits to client sites; participate in pre-sales meetings&lt;br /&gt;
&lt;br /&gt;
* Manage technical aspects of client onboarding&lt;br /&gt;
&lt;br /&gt;
* Generate and prototype ideas for new products and features&lt;br /&gt;
&lt;br /&gt;
* Manage small development teams for new initiatives&lt;br /&gt;
&lt;br /&gt;
* Work with product managers to produce development roadmaps&lt;br /&gt;
&lt;br /&gt;
* Coordinate with development teams in Beijing&lt;br /&gt;
&lt;br /&gt;
* Work with software teams to understand client workflows and educate design and implementation of features and advise on user experience.&lt;br /&gt;
&lt;br /&gt;
Apply here&amp;lt;br&amp;gt;&lt;br /&gt;
https://toc.taleo.net/careersection/2/jobdetail.ftl?lang=en&amp;amp;job=TEC00022861&lt;br /&gt;
&lt;br /&gt;
==SEMANTIC WEB PROGRAMMER/DEVELOPER at Brown University Library in Providence, RI == &lt;br /&gt;
Brown University is in the process of implementing VIVO, an open source semantic web application that supports and facilitates research discovery within and among institutions. The Brown University Library is seeking a Semantic Web Programmer/Developer to play a vital role in the launch and support of this new campus enterprise system. This full time, permanent position is an exciting opportunity for a programmer with experience in semantic web technologies to advance a large-scale linked open data project. &lt;br /&gt;
&lt;br /&gt;
The Semantic Web Programmer/Developer is responsible for initial data ingest planning and execution, for configuration of local extensions to the application ontology and for ongoing maintenance of ontology and data. The position develops and documents scripts using XML and semantic web technologies to process data and metadata from institutional databases of record, online databases of publications and research grant information, and other sources as identified by campus stakeholders. S/He writes programs and web services to return integrated and enhanced data to institutional stakeholders in RDF, XML, JSON, and other formats for reporting analysis, archiving, and display. The position participates actively in the VIVO development network and represents Brown in the national VIVO community. The position reports to the Head, Integrated Technology Services in the Brown University Library. &lt;br /&gt;
&lt;br /&gt;
Qualifications: &lt;br /&gt;
* Bachelor’s Degree in Computer Science or Advanced Degree in Information Science; plus three to five years relevant experience &lt;br /&gt;
* Experience working with RDF data model and semantic web design principles &lt;br /&gt;
* Experience with formal ontology languages such as OWL and RDFS &lt;br /&gt;
* Experience with languages for querying RDF (e.g., SPARQL, SeRQL) &lt;br /&gt;
* Experience with one or more metadata manipulation and scripting languages: XSLT, Java, Perl, Python, or PHP. &lt;br /&gt;
* Knowledge of data extraction, mining, harvesting techniques and tools &lt;br /&gt;
* Excellent interpersonal, oral and written communication skills &lt;br /&gt;
* Ability to work independently and as a member of a team &lt;br /&gt;
&lt;br /&gt;
Preferred Qualifications: &lt;br /&gt;
Experience with Jena or other semantic web libraries &lt;br /&gt;
Experience with deployment of RDF triple stores in a production environment &lt;br /&gt;
Experience with metadata issues related to the discovery of academic resources &lt;br /&gt;
&lt;br /&gt;
To apply for this position (JOB #B01403), please visit Brown’s Online Employment website (https://careers.brown.edu), complete an application online, attach documents, and submit for immediate consideration. Documents should include cover letter, resume, and the names and e-mail addresses of three references. Review of applications will continue until the position is filled. Brown University is an Equal Opportunity/ Affirmative Action Employer. &lt;br /&gt;
&lt;br /&gt;
You can find the online description[http://library.brown.edu/about/employment.php#semantic here]&lt;br /&gt;
&lt;br /&gt;
== Linked Data Engineer at Financial Institution in NYC==&lt;br /&gt;
&lt;br /&gt;
Date:2/29/2012&lt;br /&gt;
&lt;br /&gt;
Our client, a leading financial institution, is looking for a Linked Data Engineer to work on high-profile projects within their engineering and architecture group. This position is a mission-critical opportunity, which will improve user experience globally for the client, working with linked data, RDF/OWL modeling, and Jena/Protege.&lt;br /&gt;
&lt;br /&gt;
On a daily basis, this individual will have opportunity to be involved in analysis, design, and integration of a firm-wide linked open data solution. This individual will be the on-site expert working on cutting edge technologies. This person will have the opportunity to be involved with high profile projects before they are released to the rest of the bank.&lt;br /&gt;
&lt;br /&gt;
To be qualified for this position, this individual should have strong knowledge of RDF/OWL modeling and either Jena, Protégé, or D2RQ. It would be a plus if this individual has exposure to J2EE technologies such as Tomcat, Spring, Eclipse, Maven, RMI, Java Server Faces, and JMS.&lt;br /&gt;
&lt;br /&gt;
This position is long-term with the opportunity for growth. This opportunity offers the opportunity to learn new skills and be part of a cutting edge team. Please reply if you are available for an interview within 72 hours.&lt;br /&gt;
&lt;br /&gt;
Online description: [http://information-technology.thingamajob.com/jobs/New-York/Linked-Data-Engineer/2496518 thingamajob]&lt;br /&gt;
&lt;br /&gt;
Please contact Lauren Freidhof for further information: lfreidho@teksystems.com&lt;br /&gt;
&lt;br /&gt;
== Research Engineer - The New York Times Company==&lt;br /&gt;
&lt;br /&gt;
The New York Times Company, a leading media company with 2010 revenues of $2.4 billion, includes The New York Times, the International Herald Tribune, The Boston Globe, 15 other daily newspapers and more than 50 Web sites, including NYTimes.com, Boston.com and About.com. The Company’s core purpose is to enhance society by creating, collecting and distributing high-quality news, information and entertainment.&lt;br /&gt;
&lt;br /&gt;
The New York Times Company&#039;s Research and Development group is looking for a talented developer to serve as the Research Engineer for Knowledge Management.  The New York Times has one of the world&#039;s great archives and our ideal candidate will be passionate about creating innovative prototypes, systems, standards and products from this amazing resource.  We are looking for someone with an innate curiosity and a passion for innovation and who has the ability to channel this passion into both individual and team projects.&lt;br /&gt;
 &lt;br /&gt;
Responsibilities&lt;br /&gt;
*Conceptualize and develop prototypes for innovative knowledge management products and services with an emphasis on linked data technologies&lt;br /&gt;
*Solidify existing prototypes into production products and systems&lt;br /&gt;
*Keep abreast of the evolving knowledge management landscape including the emergence of new standards, technologies and systems&lt;br /&gt;
*Deeply understand the New York Times Company&#039;s data assets and how these assets can be leveraged to create useful systems and compelling new experiences&lt;br /&gt;
 &lt;br /&gt;
Qualifications&lt;br /&gt;
*Bachelor&#039;s degree in Computer Science or related subject, Master&#039;s degree a plus&lt;br /&gt;
*3-5 years of experience in architecting, developing and delivering complex knowledge management solutions&lt;br /&gt;
*Passion for knowledge management and big data problems&lt;br /&gt;
*Flexibility to work with a variety of technologies&lt;br /&gt;
*Strong coding skills in at least one statically typed language - Java / C# / Objective C / C++&lt;br /&gt;
*Proficient coding skills in at least one dynamically typed language - Python / Ruby / etc.&lt;br /&gt;
*Experience with PHP strongly desired&lt;br /&gt;
*Proficient in at least one relational database technology - MySQL / Oracle / Access&lt;br /&gt;
*Experience or interest in learning in NoSQL database technologies - MongoDB / RDF Triplestores / etc.&lt;br /&gt;
*Familiarity with basic approaches to natural language processing&lt;br /&gt;
*Interest in the development and promotion of industry-wide digital standards&lt;br /&gt;
*Strong communication skills and experience with explaining deeply technical concepts to non-technical audiences&lt;br /&gt;
*Ability to thrive in a self-directed environment&lt;br /&gt;
 &lt;br /&gt;
The New York Times Company is an equal employment opportunity employer, and does not discriminate on the basis of race, color, religion, gender, sexual orientation, marital status, age, disability, national origin, citizenship or any other protected characteristic. The New York Times Company is committed to diversity in its most inclusive sense.&lt;br /&gt;
&lt;br /&gt;
[http://developer.nytimes.com/docs/The_Semantic_API apply here]&lt;br /&gt;
&lt;br /&gt;
==Metropolitan Musem of Art hiring a Semantic Web Developer==&lt;br /&gt;
&lt;br /&gt;
Media Lab&amp;lt;br&amp;gt;&lt;br /&gt;
Digital Media Department&amp;lt;br&amp;gt;&lt;br /&gt;
Metropolitan Museum of Art&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
the Metropolitan Museum of Art&#039;s Digital Media Department is hiring for an Information Systems Developer.&lt;br /&gt;
This position will be involved in advanced data architecture solutions, to support a variety of web and in-gallery technology.&lt;br /&gt;
&lt;br /&gt;
This work may entail:&lt;br /&gt;
- Setting up and administering triple stores, NoSQL dbs, and CMSs like Drupal&lt;br /&gt;
- designing interfaces, modules, and workflows for same&lt;br /&gt;
- Implementing collective intelligence algorithms, &lt;br /&gt;
- experimenting with new technologies, developing prototypes and proofs-of-concept&lt;br /&gt;
- and (to be honest) some drudgery, like data delivery, ETL, and report generation&lt;br /&gt;
&lt;br /&gt;
See the application on linkedin [http://www.linkedin.com/jobs?viewJob=&amp;amp;jobId=2157751&amp;amp;srchIndex=0&amp;amp;trk=njsrch_hits&amp;amp;goback=%2Efjs_information+systems+developer_*1_*1_I_us_*1_*1_1_R_true_*2_*2_*2_*2_*2_*2_*2_*2   here].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
I know many of you do more than just SemWeb work, and many of you are on this list because you like to find new ways to tackle vexing problems. That&#039;s what we&#039;re looking for.&lt;br /&gt;
&lt;br /&gt;
If you choose to submit a resume, please send it to the email address provided, but also cc: don.undeen@metmuseum.org&lt;br /&gt;
&lt;br /&gt;
== Job/Contracting Work - NYC ==&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 5/14/2011&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Hello,&lt;br /&gt;
 &lt;br /&gt;
I am looking to speak to people with Natural Language Processing/Machine Learning experience regarding a potential project to help analyze Social Media data/perform sentiment analysis.  Ideal candidate will have an academic background in different NLP/Machine Learning theories, frameworks and real-world implementations.  A background in programming is a plus.  Pay is competitive.  Please send me a message if you are interested in having a discussion to learn more about the opportunity.  Please contact me at: igorgonta@yahoo.com.&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Gluejar is Hiring ==&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039;3/1/2011&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
We have funding for 4 positions, ranging from Engineering to Marketing, in a metadata-rich problem space.&lt;br /&gt;
&lt;br /&gt;
For details, please see [http://go-to-hellman.blogspot.com/2011/03/gluejar-is-hiring.html this post on the Go To Hellman blog].&lt;br /&gt;
&lt;br /&gt;
== Metadata Manager at the Associated Press (New York City) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039;2/25/2011&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Associated Press is seeking a Metadata Manager for its New York City and/or Cranbury, NJ office locations.&lt;br /&gt;
&lt;br /&gt;
The Metadata Manager will be responsible for developing and maintaining metadata standards across AP’s diverse set of content and products. Responsibilities will include all aspects of working with metadata schema, such as information modeling, XML validation, transformation and testing. Additional duties will include maintaining schema artifacts (like XML schema definitions, metadata mapping tables and editorial workflow documentation), performing metadata and content analysis, and developing technical specifications for populating and transforming metadata structures and values.  This position reports to the Deputy Director of Schema Standards and is within the AP’s Information Management group.&lt;br /&gt;
&lt;br /&gt;
Primary responsibilities&lt;br /&gt;
&lt;br /&gt;
* Collaborate with journalists, technologists, product development and members of the Information Management team to define and document metadata schema across media types and products.&lt;br /&gt;
&lt;br /&gt;
* Gather business requirements for content and content metadata, and capture and maintain business rules related to metadata modeling and transforms.&lt;br /&gt;
&lt;br /&gt;
* Design, develop and execute content and metadata analyses, interpret results and provide written summaries and reports.&lt;br /&gt;
&lt;br /&gt;
* Work with Development and QA to specify and test schemas and transforms and to communicate changes to stakeholders.&lt;br /&gt;
&lt;br /&gt;
Knowledge, Skills and Abilities&lt;br /&gt;
&lt;br /&gt;
* Familiarity with XML and XML schema languages. Knowledge of XPath, XSLT or XQuery a plus.&lt;br /&gt;
&lt;br /&gt;
* Experience creating and implementing metadata schemas and taxonomies or controlled vocabularies.&lt;br /&gt;
&lt;br /&gt;
* Experience providing requirements to technical teams. Hands on programming experience a plus.&lt;br /&gt;
&lt;br /&gt;
* Familiarity with databases and query languages. Statistical data or content analysis experience a plus.&lt;br /&gt;
&lt;br /&gt;
* Familiarity with the publishing, entertainment or media industries a plus.&lt;br /&gt;
&lt;br /&gt;
* Excellent written and oral communications skills.&lt;br /&gt;
&lt;br /&gt;
* Demonstrated ability to work effectively across groups to achieve objectives.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039;&lt;br /&gt;
To learn more about the position and to apply:&lt;br /&gt;
https://careers.ap.org/viewjob.html?optlink-view=view-19563&amp;amp;ERFormID=newjoblist&amp;amp;ERFormCode=any&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==University of Minnesota Libraries Seeks Drupal Developers==&lt;br /&gt;
Date: 11/2/2010&lt;br /&gt;
&lt;br /&gt;
The University of Minnesota Libraries seeks two or more talented Drupal&lt;br /&gt;
software developers, for either one or two year appointments, to design and&lt;br /&gt;
support new, innovative web-based library services, systems, and tools which&lt;br /&gt;
address as well as anticipate the evolving needs of library users.&lt;br /&gt;
&lt;br /&gt;
The University Libraries are supporting multiple projects using the Drupal&lt;br /&gt;
platform.  Responsibilities could include two or more of the following areas&lt;br /&gt;
of Drupal development:&lt;br /&gt;
&lt;br /&gt;
*In collaboration with our partners in the American Indian Studies department, provide primary development support for the forthcoming Online Ojibwe Dictionary. Responsibilities include addressing both content provider and user needs in developing a robust web application using the Drupal framework.&lt;br /&gt;
*Provide development support for the University Libraries&#039; UMedia Archive (umedia.lib.umn.edu), a digital library application that provides users with access to many of the Libraries rich media collections as well as allowing for user submitted uploads. Using Drupal, work to integrate new features and support current mechanisms that help further the enhance the user experience.&lt;br /&gt;
*Assist in implementing the Drupal CMS for the main public facing web site of the University Libraries (www.lib.umn.edu) , creating customization and personalization options for library users, helping in the creation of mobile version of library web site(s), designing new sites, and using new web services technologies to improve the user experience in discovering, searching, finding, or acquiring library materials and content. Projects may also likely include further integration of library resources into the Moodle course management tool, implementation of Shibboleth identity management system, and creatively using various API&#039;s made available by Google, OCLC, Amazon, Ex Libris and other library vendors.&lt;br /&gt;
&lt;br /&gt;
For more information and to apply:&lt;br /&gt;
&lt;br /&gt;
http://employment.umn.edu/applicants/Central?quickFind=90989&lt;br /&gt;
&lt;br /&gt;
== Early Stage, Funded Start Up with $2B Pilot Customer Seeks Lead Engineer/ VP Development/Tech Assassin - CovetedList.com ==&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039;8/8/10&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Company:          CovetedList.com &lt;br /&gt;
&lt;br /&gt;
Elevator Pitch:   We provide smart data personalization services when shopping, browsing and sharing on the internet. [Translation: Our goal is to be the next generation of Nielsens or NPD Group using a combination of linked data technologies, rule based information retrieval with natural language processing and learning systems/ machine learning]&lt;br /&gt;
&lt;br /&gt;
i.e. we have a REALLY cool idea and a very disruptive business model. &lt;br /&gt;
&lt;br /&gt;
411:              Early stage, *funded* start up (with a Fortune 1000 pilot customer), looking for the Head of Development who can build a team and a platform and has a good understanding of natural language processing, data extraction, information retrieval, machine learning and familiar with semantic technologies, of course.  The ideal candidate has a MS or PhD or PhD candidate, is an out of the box academic thinker who knows how to code in Python/C++.&lt;br /&gt;
&lt;br /&gt;
We are planning to build our front end using modern, web-facing technologies (JavaScript, HTML5, etc.), and implement our back-end services and algorithms in the most dynamic programming environment that is up to the task (but momentum as of late is *strongly* leaning towards Python). The candidate will solidify this decision for us. (Note: our core libaries and ontologies are pretty built out, we are building apps for them to become living entities- which would be part of your job)&lt;br /&gt;
&lt;br /&gt;
Candidates should have experience in the product management cycle, on-time delivery, and great communication skills. Agile and XP development methodologies are good skills. Scrum is cool. If you don&#039;t know something, it is important to admit it and be psyched to learn.&lt;br /&gt;
&lt;br /&gt;
MUST BE LOCATED IN OR AROUND THE NY METRO AREA.  WE ARE IN NYC.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; charlotte@covetedlist.com&lt;br /&gt;
&lt;br /&gt;
== User Interface Developer - OrangeDog ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 5/6/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
I am putting together a &amp;quot;semantic applications&amp;quot; startup, with several semantic applications in mind.&lt;br /&gt;
&lt;br /&gt;
The first area I am looking at is data integration and user interfaces using ontologies. I have already written software for both the data integration and the user interface side of things. The software works quite nicely -- it is well beyond a prototype (it&#039;s stable, modular, not too slow, etc). The integration side of the software is much more mature than the user interface side (I am no user interface designer!). &lt;br /&gt;
&lt;br /&gt;
I have also written a detailed business plan. I am now actively looking for startup funding (the usual angel, VC, etc stuff), and I am looking for people who may be interested in joining the startup, as both employees, and as equity holders, founders, etc.  &lt;br /&gt;
&lt;br /&gt;
I am specifically interested in finding someone with interest and expertise in user interfaces. &lt;br /&gt;
&lt;br /&gt;
This isn&#039;t a user experience position, it&#039;s an engineering/coding position. The UI person will have two roles. The first is the standard one of writing UIs for the various applications we have. Pretty straightforward stuff. The second role is to build so called &amp;quot;semantic&amp;quot; user interfaces. These are interfaces that are generated from ontologies. This role is very open ended. There is basically an endless scope here for a creative user interface developer, as there are a lot of interesting things you can do when generating UIs from ontologies.&lt;br /&gt;
&lt;br /&gt;
Essentially the person I am looking for needs to be a really good UI coder, plus able to understand what ontologies bring to the table (you don&#039;t necessarily need to know this now, but be capable of learning it), and hence how users may want to interact with them visually, plus have a good visual sense, since the UI person will be in charge of the user interface side of our product suite. They must have a passion for understanding and learning how end users want to use both user interfaces and ontologies, and hence how they can support those end users -- we are relentlessly client focussed.&lt;br /&gt;
&lt;br /&gt;
To be clear -- I cannot pay anyone at this point as I haven&#039;t got funding yet. So for the moment this all has to be for equity consideration. If we do get funded this turns into a full time job.&lt;br /&gt;
&lt;br /&gt;
We are located on the west coast of the US in Los Angeles, but you don&#039;t have to be. Getting quality people is more important to us than having them in the &amp;quot;right&amp;quot; location. If you are interested, or know anyone interested, please drop me a line at graham@orangedogconsulting.com.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Web Developer - Reflexions Data ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: White Plains (Westchester/NYC)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 4/2/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Reflexions Data is a growing company of creative innovators that develops web applications for clients in a variety of industries including publishing, marketing, retail e-commerce, internet startups and non-profit organizations. We are an equal opportunity employer (M/F/D/V).&lt;br /&gt;
&lt;br /&gt;
We&#039;ve been around for 10+ years and our office is a casual, fast-paced environment.  We offer competitive compensation and benefits including group health insurance and a 401(k) retirement plan.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
* Must be willing to work in a closely knit team environment and must demonstrate a passion for solving business problems with technology.&lt;br /&gt;
* BA/BS or MS in Computer Science or related technical discipline.&lt;br /&gt;
* Deep understanding of computer science fundamentals, including data structures, algorithms, and software design principles.&lt;br /&gt;
* Familiarity with MVC-style web development frameworks.&lt;br /&gt;
* At least 2 years of web/software application design and development experience.&lt;br /&gt;
* Extensive knowledge and experience with UNIX/Linux.&lt;br /&gt;
* Intimate familiarity with web standards and front-end technologies including XHTML, CSS, and Javascript/AJAX.&lt;br /&gt;
&lt;br /&gt;
Do you read Slashdot every day? Do you see regular expressions in your dreams or write Python code for fun? Then we encourage you to introduce yourself! This is a great mid-level position with opportunities for advancement. &lt;br /&gt;
&lt;br /&gt;
Learn more and apply here:&lt;br /&gt;
http://www.reflexionsdata.com/company/employment&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Metadata Librarian/Analyst - Bloomberg ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: Skillman, NJ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 4/1/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Company&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Bloomberg is the world’s most trusted source of information for businesses and professionals. Bloomberg combines innovative technology with unmatched analytic, data, news, display and distribution capabilities, to deliver critical information via the BLOOMBERG PROFESSIONAL® service and multimedia platforms. Bloomberg&#039;s media services cover the world with more than 2,200 news and multimedia professionals at 146 bureaus in 72 countries. The BLOOMBERG TELEVISION® 24-hour network delivers smart television to more than 240 million homes. BLOOMBERG RADIO® services broadcast via SIRIUS XM Radio and 1worldspaceTM satellite radio globally and on WBBR 1130AM in New York. The award-winning monthly BLOOMBERG MARKETS® magazine, Bloomberg BusinessWeek magazine and the BLOOMBERG.COM® financial news and information Web site provide news and insight to businesses and investors.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Role&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Bloomberg is looking for a highly motivated individual to join a newly formed information-retrieval group.  As a Librarian/Analyst within this group, you will be part of a team dedicated to helping provide the next generation in Bloomberg Terminal usability to our clients by standardizing, organizing, and ensuring accuracy of keyword databases to support a new information-retrieval system. The successful candidate will apply taxonomy and ontology principles to create and manage metadata to improve ‘findability’ of resources, diagnose and fix inconsistencies, solve reference issues, analyze linguistic usage patterns and perform basic data analysis tasks.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;UL&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt; Bachelor&#039;s degree or equivalent work experience.  Masters in Library Science (MLS, MLIS), Bachelor of Science in Computer Science, or related technical discipline is preferred. &lt;br /&gt;
&amp;lt;li&amp;gt; 3 or more years of data analysis experience.&lt;br /&gt;
&amp;lt;li&amp;gt; Experience in use of relational databases, such as Access, MySQL, or Oracle. &lt;br /&gt;
&amp;lt;li&amp;gt; Knowledge of metadata standards such MARC, Dublin Core, etc., and the proven ability to apply metadata to large scale collections of data.&lt;br /&gt;
&amp;lt;li&amp;gt; Demonstrated ability to research and analyze problems and develop solutions.&lt;br /&gt;
&amp;lt;li&amp;gt; Strong analytic and organization skills.&lt;br /&gt;
&amp;lt;li&amp;gt; Ability to exercise independent judgment.&lt;br /&gt;
&amp;lt;li&amp;gt; Experience handling financial, economic and discrete content is helpful, but not required.&lt;br /&gt;
&amp;lt;li&amp;gt; Data modeling experience is a plus.&lt;br /&gt;
&amp;lt;/UL&amp;gt;&lt;br /&gt;
Bloomberg is an equal opportunity/affirmative action employer and we welcome applications from all backgrounds regardless of race, color, religion, sex, national origin, ancestry, age, marital status, sexual orientation, gender identity, veteran status, disability, or any other classification protected by law.&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
Please apply online at http://careers.bloomberg.com/hire/jobs/job25667.html&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Senior Software Engineer - daylife==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: New York City&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 3/18/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Come to Daylife and help us build the future of online news.  As a Senior Software Engineer, you&#039;ll develop scalable systems for organizing and delivering information to people and organizations around the world.  Add new features to our customer APIs; improve the performance of our search systems; enhance the quality of our information extraction; evolve our engineering tools to tighten our release cycles.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
* Strong problem solving spirit.&lt;br /&gt;
* Expert in (C || Java) &amp;amp;&amp;amp; Python.&lt;br /&gt;
* Expert UNIX network programming skills -- IP protocols, sockets, IPC, event-driven programming and frameworks (libevent, NIO).&lt;br /&gt;
* Expert knowledge of Linux programming and POSIX operating system concepts, and debugging in a Linux environment -- processes, pthreads, memory model, filesystems, system introspection and debugging tools.&lt;br /&gt;
* Strong software design skills.  Write well-organized, maintainable and testable code.&lt;br /&gt;
* Solid knowledge of version control with subversion; git knowledge a plus.&lt;br /&gt;
* RHEL / CentOS administration or operating experience a plus.&lt;br /&gt;
* Strong leadership, teamwork, and communication skills.&lt;br /&gt;
* Strong coaching skills; proactively shares knowledge.&lt;br /&gt;
&lt;br /&gt;
Please contact with your CV Ken Ellis (ken [at] daylife.com)&lt;br /&gt;
&lt;br /&gt;
http://www.daylife.com/&lt;br /&gt;
&lt;br /&gt;
== Senior Semantic and Search Architect - Financial Services==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: New York City&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 3/15/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Responsible for architecture of a content mining and semantic based product offerings for the Knowledge Management Practice Area. Researches, analyzes, and recommends technologies to develop new and enhance existing systems to support&lt;br /&gt;
semantic requirements. Serves as a technical expert and is responsible for resolving complex problems and guiding development to improve the application of search and semantic technologies.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Required Skills and Experience:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
BS in Computer Science or a related field, MS preferred&amp;lt;br&amp;gt;&lt;br /&gt;
Extensive experience with distributed software development&amp;lt;br&amp;gt;&lt;br /&gt;
Data mining and predictive modeling skills&amp;lt;br&amp;gt;&lt;br /&gt;
Strong OO knowledge&amp;lt;br&amp;gt;&lt;br /&gt;
Expertise in C#.Net or Java&amp;lt;br&amp;gt;&lt;br /&gt;
Comfortable with distributed development environments and tools&amp;lt;br&amp;gt;&lt;br /&gt;
Minimum of 10 years development experience&amp;lt;br&amp;gt;&lt;br /&gt;
Excellent communication skills&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with Search technologies – Autonomy IDOL/ FAST/ Verity / Lucene&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with mentoring software development resources.&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with scoping the technical side of client facing products that require scale and contextual consumer experiences.&lt;br /&gt;
Experience with understanding usage flows and has experience with mining historical consumer choice data to make a service more intelligent about consumer options.&amp;lt;br&amp;gt;&lt;br /&gt;
Experience in technical architectures for search applications, including data structures and algorithms to support entity extraction, disambiguation, normalization and clustering.&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with architecture models and design patterns to support development and maintenance of taxonomies and knowledge bases. Ideal candidate has been working on ways to accomplish this in an enterprise environment using open source&lt;br /&gt;
or by using data available from multiple sources.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills Desired&#039;&#039;&#039;&amp;lt;BR&amp;gt;&lt;br /&gt;
Background in natural language processing NLP&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with cloud computing platforms&amp;lt;br&amp;gt;&lt;br /&gt;
Exposure to popular open source machine learning tools&amp;lt;br&amp;gt;&lt;br /&gt;
Data harvesting experience&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with search content classification&amp;lt;br&amp;gt;&lt;br /&gt;
Experience in the publishing industry particularly in optimization&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with non-relational NoSQL document oriented data stores&amp;lt;br&amp;gt;&lt;br /&gt;
Understanding of Semantic Web nomenclature and related technologies&amp;lt;br&amp;gt;&lt;br /&gt;
Key words: semantic web, Resource Description Framework RDF, Autonomy IDOL,Verity K2, data interchange formats, RDF, XML, N3, Turtle, N-Triples, RDF Schema RDFS and the Web Ontology Language OWL, semantic web stack, URI, Ontologies, NoSQL&lt;br /&gt;
&lt;br /&gt;
Please contact with your attached CV: info@kona.llc&lt;br /&gt;
&lt;br /&gt;
== User Interface Designer - Kikin==&lt;br /&gt;
&lt;br /&gt;
Kikin is searching for a User Interface Designer to work directly with our VP of Product on the user interface design. You will also be responsible for creating working product mock-ups for our key clients and partners that demonstrate Kikin&#039;s custom product capabilities.&lt;br /&gt;
&lt;br /&gt;
Our user interface designer will help us fulfill our mission of making it SIMPLE to navigate the web. This person will be the leader when it comes to making this process much easier than it is today. You will be asked to create beautiful designs that are obvious to our users. This isn&#039;t going to be a run of the mill HTML and CSS job - we&#039;re looking for someone that is going to dig in and create an experience that makes our web platform legendary. You&#039;ll be running the UI show which will include: brainstorming design and usability options, building simple wire frames, deliver clean, elegant HTML, CSS and javascript. We have ideas for how our features might look good--you know how to design them so we get unsolicited calls about how amazing the site is.&lt;br /&gt;
&lt;br /&gt;
You get a thrill from building things people love to use. You think the best interfaces are the ones that get out of the way and let people do their work, not necessarily the ones that grab the most attention. You&#039;re constantly finding the sweet spot between beauty and usability. You know users are impatient and you don&#039;t have much time to impress them. You&#039;re oozing with creativity. When someone asks you how you&#039;d design something, you immediately think of ten completely different options. You can mock them up in minutes, not hours, and you can objectively evaluate the most usable. You&#039;re as comfortable drawing sketches on paper napkins as you are throwing layers around in Adobe Creative Suite 4. You feed off the energy and enthusiasm of others. Simply put, you love the challenge of building a great user experience from inception to delivery.&lt;br /&gt;
What you need:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* 2+ years experience with HTML/CSS development and graphic design for professional web projects&lt;br /&gt;
* Detailed knowledge of HTML/CSS standards &lt;br /&gt;
* Basic javascript proficiency &lt;br /&gt;
* Graphic Design (You really know your way around a professional graphics package)&lt;br /&gt;
* Strong understanding of web application usability&lt;br /&gt;
* Portfolio of some of your past work&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;About Us&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
At Kikin, we’re driven to revolutionize online &amp;amp; mobile navigation, services, and advertising.  Using proprietary client-side technology, kikin provides on-the-fly, browser-based enhancements that deliver personalization, increased relevance, richer merchandising and content, and integrated services-- all without users having to change any of their existing behavior.  kikin supports all popular search engines, major service providers, browsers, operating systems, etc.&lt;br /&gt;
&lt;br /&gt;
Kikin offers a fantastic environment and a team of bright, dynamic people from all over the world. We are dedicated to fundamentally changing the way people navigate and use the Internet and having a great time doing it.&lt;br /&gt;
&lt;br /&gt;
While maintaining a low profile, we’re testing in beta and continue to win key partnerships with major content, commerce, service providers, OEM distributors and developers. Kikin was founded by seasoned entrepreneurs with successful track records.  Operations have been established in the U.S. (NYC) and Europe (Berlin), and will be established in China (Shanghai) and Japan (Tokyo) by the end of Q4CY09. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
This position is full-time in New York City (SOHO).  This is not a telecommute or contract role.&lt;br /&gt;
No third-party, sub-contractors/agencies. Unfortunately sponsorship are Not available.&lt;br /&gt;
&lt;br /&gt;
We will require credible work references/background check.&lt;br /&gt;
Salary commensurate with experience.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Team Lead/ Back End Search, Data Engineer - kikin==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The Java Developer position at kikin is working on enhancing kikin’s next-generation internet application and services. &lt;br /&gt;
Software development tasks are focused on information retrieval, data mining, relationship mapping and collaborative filtering.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Responsibilities&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Lead a small team that will:&lt;br /&gt;
* Participate in all stages of development including Design, Implementation, and Testing of our next-generation Federated search and Personalization back end&lt;br /&gt;
* Integrate Content, Commerce,  or Service Partner data and functionality into our backend&lt;br /&gt;
* Collaborate with analysts and product management on engaging features and algorithms&lt;br /&gt;
* Work on specifications that address evolving business requirements, user interfaces, process flow, performance, and scalability&lt;br /&gt;
* Reports to CTO&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills required&#039;&#039;&#039;&lt;br /&gt;
* 4+ years of Advanced Java; Expertise in performance-oriented Java&lt;br /&gt;
* Service Oriented Architecture (SOA) experience is paramount&lt;br /&gt;
* SOAP, REST, Caching, Clustering, Distributed Computing, XML Object Mapping, Web Services&lt;br /&gt;
* Experience with Wicket, Spring, AJAX, DHTML, JavaScript, jQuery, CSS&lt;br /&gt;
* Expert XML, MySQL, Apache, Resin, Linux&lt;br /&gt;
* Experience with Compass, Lucene, SOLR desired&lt;br /&gt;
* High throughput production experience preferred&lt;br /&gt;
* Test-driven development and iterative self-correction must be second nature&lt;br /&gt;
* Self-directed, highly motivated, and able to work in a fast paced startup environment&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Database Architect - Edifice==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Northern New Jersey&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; February 20, 2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity:&#039;&#039;&#039; Our innovative company is looking for a brilliant database architect to design the back-end of a new SaaS product.  If this is you or someone you know, please contact me directly:  nbruce[at]EdificeInfo.com&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Senior Software Engineer, machine learning &amp;amp; distributed computing experience, Java, OO, Linux ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York City&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; January 30, 2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contract&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We are a consulting and recruiting firm. Our client, an innovative advertising startup, is seeking a creative Senior Software Engineer with machine learning and distributed computing experience. This is a hands on position that requires extensive development for release in a production environment. The ideal candidate will have good analytical and troubleshooting skills, fluency in coding and excellent communication skills.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Responsibilities&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Work with an open/friendly team of engineers and research scientists&lt;br /&gt;
* Contribute to product vision/direction&lt;br /&gt;
* Create robust high-transaction production applications&lt;br /&gt;
* Develop prototypes for research projects&lt;br /&gt;
* Production application development based on research&lt;br /&gt;
* Production software and environment troubleshooting&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* BS in Computer Science or a related field, MS preferred&lt;br /&gt;
* Extensive experience with distributed software development&lt;br /&gt;
* Data mining and predictive modeling skills&lt;br /&gt;
* Strong OO knowledge&lt;br /&gt;
* Expert in Java&lt;br /&gt;
* Comfortable with distributed development environments and tools&lt;br /&gt;
* Expert in the Linux/OS X command line&lt;br /&gt;
* Proven list of shipped products&lt;br /&gt;
* Solid math background&lt;br /&gt;
* Excellent communication skills&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Desired&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Background in natural language processing NLP&lt;br /&gt;
* Previous startup experience&lt;br /&gt;
* Experience with cloud computing platforms&lt;br /&gt;
* Exposure to popular open source machine learning tools&lt;br /&gt;
* Data harvesting experience&lt;br /&gt;
* Experience with search content classification or spam detection&lt;br /&gt;
* Experience in the advertising industry particularly in optimization&lt;br /&gt;
* Experience with non-relational NoSQL document oriented data stores&lt;br /&gt;
* Understanding of Semantic Web nomenclature and related technologies&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Critical Key words: &#039;&#039;&#039; semantic web, Resource Description Framework RDF, a variety of data interchange formats, RDF,XML, N3, Turtle, N-Triples, RDF Schema RDFS and the Web Ontology Language OWL,semantic html, Resource Description Framework RDF, Web Ontology Language OWL, Extensible Markup Language XML. RDF, OWL, and XML, semantic web stack, uri, rdfs, ontologies, NoSQL, Friend of a Friend or FoaF, DBpedia&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;good to have key words:&#039;&#039;&#039; trust metric, Friendster, livejournal, virtual communities, Slashdot, karma, rummble.com, transitivity, advogato, reputation systems, subjective logic, applied computational trust, trust management, trust engines, risk engines, trustworthy recommenders, trust-based collaborative filtering, reputation and recommendation, evidence gathering, security through collaboration, technical trust, user trust, network of trust, web of trust, impact of social networking on trust and security, virtual and self-organization, decentralized identity management, trust metrics analysis, aggregation analytics, marketing metrics, social web, viral web, Trust management, Trustos, Nuglets in mobile ad-hoc networks, Slashdot.org&#039;s Karma, Ebay&#039;s feedback rating, FOAF trust module, Free Haven direct and meta trust; direct trust, Advogato&#039;s trust metric, probabilistic trust, System trust, interpersonal and self trust&lt;br /&gt;
&lt;br /&gt;
Please send resume to molly [at] fremontconsulting.com&lt;br /&gt;
&lt;br /&gt;
for additional opportunities: visit us on Facebook:&lt;br /&gt;
&lt;br /&gt;
http://www.facebook.com/business/dashboard/#/pages/Elk-Grove-CA/Fremont-Consulting/321998995452&lt;br /&gt;
&lt;br /&gt;
* Compensation: hourly&lt;br /&gt;
* This is a contract job.&lt;br /&gt;
* OK for recruiters to contact this job poster.&lt;br /&gt;
* Please, no phone calls about this job!&lt;br /&gt;
* Please do not contact job poster about other services, products or commercial interests.&lt;br /&gt;
&lt;br /&gt;
== Python Software Engineer  ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Company&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
TLists builds search and content management tools that enable leading media companies and the mass-market to take full advantage of Twitter Lists as a brand building content distribution channel.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We&#039;re looking for an intensely talented python engineer to work with our CTO (Stanford Ph.D.) and development crew to contribute to our search platform and API. We like candidates with broad skills, but to stand out you should have an exceptional record at solving hard analytic problems, and building innovative web or data mining applications.&lt;br /&gt;
&lt;br /&gt;
The ideal engineer:&lt;br /&gt;
- Has a BA or MS in computer science or a related field (E.g., cognitive science, math, statistics, or linguistics, etc.)&lt;br /&gt;
&lt;br /&gt;
- Is experienced with python, and using python for web engineering projects (e.g., search, data aggregation, data APIs, python integration with lucene/hadoop, systems architecture, AWS, etc)&lt;br /&gt;
&lt;br /&gt;
- Has interest and/or experience in some area of data mining (e.g., machine learning, computational linguistics, statistics, etc.)&lt;br /&gt;
&lt;br /&gt;
- Is based in the New York City Area, or willing to relocate&lt;br /&gt;
&lt;br /&gt;
The position comes with equity, competitive salary, health benefits, a great comp&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact Information:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Please contact marguerite (at) tlists (dot) com to learn more about this opportunity.&lt;br /&gt;
&lt;br /&gt;
== Python Web Developer  ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Location:&#039;&#039;&#039; New York, NY USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Company&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
TLists builds search and content management tools that enable leading media companies and the mass-market to take full advantage of Twitter Lists as a brand building content distribution channel.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We&#039;re looking for an intensely talented web developer to join our CTO (Stanford Ph.D.) and development crew in New York City.&lt;br /&gt;
&lt;br /&gt;
We like candidates with broad skills, but to stand out you should have strong coding skills, a keen aesthetic sense, and an exceptional record at building innovative and beautiful web sites.&lt;br /&gt;
&lt;br /&gt;
Desirable experience includes:&lt;br /&gt;
- BA or MS in a technical field (E.g., computer science, math, cognitive science, etc.)&lt;br /&gt;
&lt;br /&gt;
- Python (or similar interpreted languages)&lt;br /&gt;
&lt;br /&gt;
- Django (or similar MVC frameworks)&lt;br /&gt;
&lt;br /&gt;
- Javascript and CSS coding (e.g., jquery, YUI, prototype, etc.)&lt;br /&gt;
&lt;br /&gt;
The position comes with equity, competitive salary, health benefits, a great computing setup, and flexible working conditions.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact Information&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Please contact marguerite (at) tlists (dot) com to learn more about this opportunity.&lt;br /&gt;
&lt;br /&gt;
==  VP of Engineering / Lead Python Software Engineer ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Company&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
TLists builds search and content management tools that enable leading media companies and the mass-market to take full advantage of Twitter Lists as a brand building content distribution channel.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We&#039;re looking for a talented and experienced engineer to work alongside our CTO (Stanford Ph.D.) to lead the day-to-day efforts of the engineering team (search/API and web development groups).&lt;br /&gt;
&lt;br /&gt;
The ideal engineer:&lt;br /&gt;
- Has an M.S. or Ph.D. degree in CS or a related technical field (E.g., cognitive science, math, statistics, linguistics, etc.) and an exceptional record at solving challenging problems.&lt;br /&gt;
&lt;br /&gt;
- Has had a minimum of 5 years of professional coding experience, ideally with top search / advertising tech / or data mining oriented companies, as well as experience managing others in a team.&lt;br /&gt;
&lt;br /&gt;
- Is experienced with some or all of the following: data mining and statistics, systems architecture, search technology, building large/scalable web services, hadoop, lucene/solr, AWS, Twitter API, Twisted framework.&lt;br /&gt;
&lt;br /&gt;
- Is deeply proficient in python (Django exp a plus), and is quick to learn new languages and technologies (Java and Javascript also pluses).&lt;br /&gt;
&lt;br /&gt;
- Is based in the New York City Area, or willing to relocate.&lt;br /&gt;
&lt;br /&gt;
This position comes with significant equity, competitive salary, health benefits, a great computing setup, and flexible working conditions.&lt;br /&gt;
&lt;br /&gt;
TLists favors a relatively non-hierarchical working environment. As lead, you would spend roughly half your time developing code and the other half mentoring and coordinating the others on the team.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact Information&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Please contact marguerite (at) tlists (dot) com to learn more about this opportunity.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Senior Applications Developer-NYPL==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; October 8, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
http://jobs-nypl.icims.com/jobs/5655/job&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;General Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Under the direction of the Managing Director, NYPL Digital Labs:&lt;br /&gt;
&lt;br /&gt;
* Supports the implementation of a multi-instance Fedora repository&lt;br /&gt;
* Codes, integrates, and maintains services and applications that support digital object ingest, preservation, search, discovery, distribution, and &lt;br /&gt;
* Designs, implements, tests, and writes documentation of custom software applications&lt;br /&gt;
* Integrates and extends various open-source solutions&lt;br /&gt;
* Manipulates large metadata sets&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Eligibility Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Bachelor&#039;s degree in Computer Science or relevant field; advanced degree preferred&lt;br /&gt;
* 3-5 years of related experience.&lt;br /&gt;
* Extensive experience with relational databases, database design, and fluency in SQL&lt;br /&gt;
* Strong Java skills and object-oriented design experience, including knowledge of core libraries, servlets, JDBC Experience with Apache Web server and Tomcat application server&lt;br /&gt;
* Knowledge of RESTful architectures and HTTP; familiarity with RDF, OWL and triplestores preferred&lt;br /&gt;
* Experience with other programming languages, such as PHP, Python, or Ruby preferred&lt;br /&gt;
* Experience with Web services technologies (SOAP/WSDL) preferred&lt;br /&gt;
* Experience with the Fedora repository software is preferred&lt;br /&gt;
&lt;br /&gt;
==4 OPEN Ph.D. POSITIONS at Hasso-Plattner-Institute (HPI)==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; September 25, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Berlin, Germany&lt;br /&gt;
&lt;br /&gt;
We offer 4 OPEN Ph.D. POSITIONS at Hasso-Plattner-Institute (HPI), Potsdam (Germany) starting on the fourth quarter of 2009&lt;br /&gt;
&lt;br /&gt;
Hasso-Plattner-Institute (HPI) is a privately financed institute&lt;br /&gt;
affiliated with the University of Potsdam, Germany.&lt;br /&gt;
The Institute&#039;s founder and benefactor Professor Hasso Plattner,&lt;br /&gt;
who is also co-founder and chairman of the supervisory board of SAP AG,&lt;br /&gt;
has created an opportunity for students to experience a unique education in IT systems engineering&lt;br /&gt;
in a professional research environment with a strong practice orientation.&lt;br /&gt;
(for more information on HPI, c.f. http://www.hpi.uni-potsdam.de/ )&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Project Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
MEDIAGLOBE is part of the THESEUS research program,&lt;br /&gt;
initiated by the German Federal Ministry of Economy and Technology (BMWi),&lt;br /&gt;
with the goal of developing a new Internet-based infrastructure&lt;br /&gt;
in order to better use and utilize the knowledge available on the Internet.&lt;br /&gt;
The focus of the research program is on semantic technologies,&lt;br /&gt;
which determine contents (words, images, sounds, and videos)&lt;br /&gt;
not through conventional methods (e.g., combinations of letters)&lt;br /&gt;
but which are able to recognize and place the meaning of a content in its proper context.&lt;br /&gt;
MEDIAGLOBE deals with digitalization, analysis, and semantic retrieval&lt;br /&gt;
of historical, documentary audiovisual content.&lt;br /&gt;
(for more information on MEDIAGLOBE, c.f. http://theseus-programm.de/theseus-mittelstand-2009/ )&lt;br /&gt;
&lt;br /&gt;
The ideal candidate holds a MS degree in Computer Science or related field&lt;br /&gt;
and is able to consider both theoretical and practical/ implementation aspects in her/his work.&lt;br /&gt;
Fluent English communication and programming skills are fundamental requirements.&lt;br /&gt;
Preferably the candidate has a background in one of the following fields:&lt;br /&gt;
&lt;br /&gt;
* semantic web technologies&lt;br /&gt;
* knowledge representations and ontology engineering&lt;br /&gt;
* audiovisual retrieval and analysis&lt;br /&gt;
* semantic search&lt;br /&gt;
* innovative web development&lt;br /&gt;
* user interface design for audiovisual content&lt;br /&gt;
&lt;br /&gt;
The position starts as soon as possible and is full-time (40h/week)&lt;br /&gt;
for the duration of the project until Oct 2011.&lt;br /&gt;
Review of applications will begin immediately and will continue until the position is filled.&lt;br /&gt;
The successful candidate will work tightly with international partners&lt;br /&gt;
and has the possibility to pursue PhD work within the scope of the project.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;How to apply:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Excellent candidates are invited to apply with:&lt;br /&gt;
&lt;br /&gt;
* Curriculum vitae and copies of degree certificates/transcripts,&lt;br /&gt;
* Writing samples/copies of relevant scientific papers (e.g. thesis,etc.),&lt;br /&gt;
* Letters of recommendation.&lt;br /&gt;
&lt;br /&gt;
Please send your application in PDF format, indicating in the subject &amp;quot;Application for PhD position&amp;quot;&lt;br /&gt;
via email or traditional mail to the following contact.&lt;br /&gt;
&lt;br /&gt;
Contact and application:&amp;lt;br&amp;gt;&lt;br /&gt;
Harald Sack&amp;lt;br&amp;gt;&lt;br /&gt;
Hasso-Plattner-Institut für Softwaresystemtechnik GmbH&amp;lt;br&amp;gt;&lt;br /&gt;
Universität Potsdam&amp;lt;br&amp;gt;&lt;br /&gt;
Prof.-Dr.-Helmert-Str. 2-3&amp;lt;br&amp;gt;&lt;br /&gt;
D-14482 Potsdam, Germany&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
phone: 	+49 (0)331-5509-527&amp;lt;br&amp;gt;&lt;br /&gt;
fax: 	+49 (0)331-5509-325&amp;lt;br&amp;gt;&lt;br /&gt;
email: 	harald.sack[AT]hpi.uni-potsdam.de&amp;lt;br&amp;gt;&lt;br /&gt;
web:   	http://www.hpi.uni-potsdam.de/meinel/persons/sack.html&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Exciting, semantic web search startup looking for Web/Data Mining Engineer==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; August 28, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Menlo Park, CA, USA&lt;br /&gt;
&lt;br /&gt;
I’m reaching out to the group on behalf of a Series B semantic web search startup in Menlo Park&lt;br /&gt;
that I’m currently working with who is looking for a Sr. Web/Data Mining Engineer&lt;br /&gt;
to join their 30-person team on a full-time basis.&lt;br /&gt;
This is a priority hire for them, and they’re looking to move quickly.&lt;br /&gt;
If you’re interested and would like more details about this opportunity,&lt;br /&gt;
feel free to reply directly to me, and I’ll get you the pertinent information accordingly.&lt;br /&gt;
&lt;br /&gt;
Thanks!&lt;br /&gt;
&lt;br /&gt;
-Donald&lt;br /&gt;
&lt;br /&gt;
35095&lt;br /&gt;
&lt;br /&gt;
Donald James&lt;br /&gt;
&lt;br /&gt;
Technical Recruiter&amp;lt;br&amp;gt;&lt;br /&gt;
Tel (408) 727-9000&amp;lt;br&amp;gt;&lt;br /&gt;
Fax (408) 716-8882&amp;lt;br&amp;gt;&lt;br /&gt;
dkj@terransys.com&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
http://www.terransys.com&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Follow Terran Systems on Twitter!&amp;lt;br&amp;gt;&lt;br /&gt;
www.twitter.com/terran_systems&amp;lt;br&amp;gt;&lt;br /&gt;
Are you Linked In? Let&#039;s link...or just view my profile:&amp;lt;br&amp;gt;&lt;br /&gt;
http://www.linkedin.com/in/dkjames&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Semantic Web Developer Java/Scala - Zurich==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Zürich, Switzerland&lt;br /&gt;
&lt;br /&gt;
Trialox.org is a startup company at the University of Zurich.&lt;br /&gt;
Our aim is to produce an open source platform for semantic applications,&lt;br /&gt;
as well as a content management system tailored to the needs&lt;br /&gt;
of international not-for-profit organizations.&lt;br /&gt;
&lt;br /&gt;
To extend our developer team,&lt;br /&gt;
we are looking for a Senior Developer experienced in programming Semantic Web applications in Java or Scala.&lt;br /&gt;
You should have a solid knowledge around Semantic Web technology&lt;br /&gt;
and ideally be experienced with the following technologies and methodologies:&lt;br /&gt;
&lt;br /&gt;
* J2SE&lt;br /&gt;
* OSGi (with Declarative Services)&lt;br /&gt;
* Maven&lt;br /&gt;
* JAX-RS&lt;br /&gt;
* Scala&lt;br /&gt;
* Scrum&lt;br /&gt;
* Test-Driven Development&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our expectation:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Apart from the technical skills,&lt;br /&gt;
we expect you to be motivated to help a small committed team to succeed.&lt;br /&gt;
This means that you know how to effectively explain your design choices,&lt;br /&gt;
as well as the occasional agreement to a compromise in order to get things done.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;What you can expect from us:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We offer a great working environment with cutting-edge technologies,&lt;br /&gt;
in a team that&#039;s committed to open source and Semantic Web standards.&lt;br /&gt;
Our ties to the University allow a continuous exchange with the latest research.&lt;br /&gt;
As a company, we are committed to maintaining a great work-life balance.&lt;br /&gt;
We offer flexible working hours as well as insurance coverage beyond the legal requirements.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;To Apply:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
If you are an excellent Software Developer and would like to apply for this position,&lt;br /&gt;
please send your CV and a cover letter demonstrating your experience to:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;hr@trialox.org&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
or&lt;br /&gt;
&lt;br /&gt;
trialox ag&amp;lt;br&amp;gt;&lt;br /&gt;
Tsuyoshi Ito&amp;lt;br&amp;gt;&lt;br /&gt;
Binzmuehlestrasse 14&amp;lt;br&amp;gt;&lt;br /&gt;
CH-8050 Zürich&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For any questions do not hesitate to contact me.&lt;br /&gt;
&lt;br /&gt;
Regards,&amp;lt;br&amp;gt;&lt;br /&gt;
Reto Bachmann&lt;br /&gt;
&lt;br /&gt;
--&lt;br /&gt;
Reto Bachmann-Gmür&amp;lt;br&amp;gt;&lt;br /&gt;
trialox.org&amp;lt;br&amp;gt;&lt;br /&gt;
Tel: +41445005015&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Senior Developer (Convention Center, in Washington, DC)==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Washington, DC, USA&lt;br /&gt;
&lt;br /&gt;
[http://washingtondc.craigslist.org/doc/sof/1260179096.html Original Posting at Craigslist]&lt;br /&gt;
&lt;br /&gt;
Semantic technologies startup in downtown Washington, DC, seeks Senior Java Developer&lt;br /&gt;
to join engineering team working on challenging problems&lt;br /&gt;
in knowledge representation, AI, and related fields.&lt;br /&gt;
&lt;br /&gt;
Qualified applicants will have&lt;br /&gt;
&lt;br /&gt;
* BS in computer science or related field; MS or PhD, ideally;&lt;br /&gt;
&lt;br /&gt;
* 5+ years of software development experience, including demonstrable Java expertise;&lt;br /&gt;
&lt;br /&gt;
* very strong software development, engineering skills;&lt;br /&gt;
&lt;br /&gt;
* background in logic, AI, KR, automated reasoning, automated planning, or related subfields, including machine learning, statistical inference; experience with Semantic Web standards, particularly OWL, ideal;&lt;br /&gt;
&lt;br /&gt;
* good verbal &amp;amp; written communication skills;&lt;br /&gt;
&lt;br /&gt;
* authorization to work full-time in the US, at our downtown DC office (telecommuting is a possibility but only in extraordinary circumstances).&lt;br /&gt;
&lt;br /&gt;
A successful applicant must be comfortable with&lt;br /&gt;
&lt;br /&gt;
* developing cutting-edge automated reasoning or automated planning systems; see Pellet (http://clarkparsia.com/pellet) or HotPlanner (http://clarkparsia.com/planner);&lt;br /&gt;
&lt;br /&gt;
* reading CS and other literature and implementing algorithms to production-quality;&lt;br /&gt;
&lt;br /&gt;
* working on bleeding-edge R&amp;amp;D projects -- i.e., think solving &amp;quot;DARPA Hard&amp;quot; problems;&lt;br /&gt;
&lt;br /&gt;
* presenting complex ideas simply to colleagues &amp;amp; customers at conferences like SemTech, ISWC, DL Workshop, OWLED, etc.;&lt;br /&gt;
&lt;br /&gt;
* familiar with open source development culture, methodologies, and process.&lt;br /&gt;
&lt;br /&gt;
About Clark &amp;amp; Parsia&lt;br /&gt;
&lt;br /&gt;
Clark &amp;amp; Parsia LLC is a bootstrap startup focusing on solving infrastructure-level problems&lt;br /&gt;
in semantic web and related areas;&lt;br /&gt;
we focus on automated reasoning, automated planning, information integration,&lt;br /&gt;
and ontology-based information systems.&lt;br /&gt;
We do advanced R&amp;amp;D in the semantic technologies area and, increasingly,&lt;br /&gt;
are focused on commercializing our research into products.&lt;br /&gt;
&lt;br /&gt;
We work in a relaxed, comfortable environment where good communication,&lt;br /&gt;
world-class coffee &amp;amp; espresso, great benefits, and competitive salaries rule the day. (Joel Test Score: 9/12)&lt;br /&gt;
&lt;br /&gt;
Clark &amp;amp; Parsia LLC is an equal opportunity employer.&lt;br /&gt;
&lt;br /&gt;
* Location: Convention Center&lt;br /&gt;
* Compensation: Commensurate with skills &amp;amp; experience&lt;br /&gt;
* Principals only. Recruiters, please don&#039;t contact this job poster.&lt;br /&gt;
* Please, no phone calls about this job!&lt;br /&gt;
* Please do not contact job poster about other services, products, or commercial interests.&lt;br /&gt;
&lt;br /&gt;
==DO YOU HAVE A BEAUTIFUL MIND?==&lt;br /&gt;
&#039;&#039;Seeking exceptional social semantic web tech start-up&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
Do you want to disrupt social networking?&lt;br /&gt;
Do you believe that someone has yet to bring a true sense&lt;br /&gt;
of purpose, relevancy and broad-scale yet efficient utility to social networking?&lt;br /&gt;
Do you have a beautiful mind that can see what is possible and then build it?  &lt;br /&gt;
&lt;br /&gt;
I was a co-founder of GoTo.com (later Overture),&lt;br /&gt;
where we revolutionized search by inventing the paid search business model.&lt;br /&gt;
Now, after a variety of other adventures&lt;br /&gt;
including a few years working on the DOE public school reform initiative in NYC&lt;br /&gt;
and some time in the Internet space in China,&lt;br /&gt;
I am developing a concept to drive social networking to the next level,&lt;br /&gt;
moving beyond the existing paradigm of reinforcing existing networks&lt;br /&gt;
to one where the network grows in new and useful ways&lt;br /&gt;
by making meaningful introductions to people one doesn’t know.  &lt;br /&gt;
&lt;br /&gt;
I am looking for an exceptional technical co-founder&lt;br /&gt;
who will provide the technical vision to my product vision -&lt;br /&gt;
someone who has extensive experience in delivering complex web-based applications of significant scale –&lt;br /&gt;
both front end and back end -&lt;br /&gt;
and who is extremely well-versed in the web technologies and standards of the “social semantic web”&lt;br /&gt;
and is capable of judiciously applying them.&lt;br /&gt;
Someone who wants to build a disruptive product&lt;br /&gt;
through the innovative application of bleeding edge web 2.0/3.0 technology.&lt;br /&gt;
And someone who is an excellent collaborator.&lt;br /&gt;
&lt;br /&gt;
You will be responsible for crystallizing the product strategy with me,&lt;br /&gt;
determining the product architecture;&lt;br /&gt;
evaluating the right platforms, database strategy and management, languages and technologies to deploy;&lt;br /&gt;
and hiring, training and leading an ace technology team to build it.&lt;br /&gt;
You should be very proficient in JAVA, AJAX, Flash, and PHP to name just a few.&lt;br /&gt;
You should embrace interoperability and data portability&lt;br /&gt;
(and the belief that users own their data)&lt;br /&gt;
and have experience with OpenID, OAuth, XFN, FOAF, XRD, XMPP, RSS, REST,&lt;br /&gt;
the Portable Contacts protocal, and the other building blocks of data portability.&lt;br /&gt;
And lastly, but importantly, you should have hands on (applied) experience in social network analysis&lt;br /&gt;
and in using machine learning, natural language and text data mining technologies&lt;br /&gt;
and other relevant semantic web tools to solve hard problems.&lt;br /&gt;
You will have a huge white canvas to work on as you will be coming in on the ground floor of this opportunity.&lt;br /&gt;
The one criterion – be passionate about building a powerful and simple platform,&lt;br /&gt;
with an elegant and intuitive UI and consumer experience.&lt;br /&gt;
&lt;br /&gt;
The position is based in NY&lt;br /&gt;
(no long-distance applications please unless you are prepared to move yourself to NYC or travel back and forth)&lt;br /&gt;
and is eligible for possible &amp;quot;CTO&amp;quot; status for the person with the right experience.&lt;br /&gt;
Ideally you are in a position to take no (or nominal) salary in favor of significant equity.  &lt;br /&gt;
&lt;br /&gt;
If this resonates with you, let’s talk.&lt;br /&gt;
I am looking for that rare person everyone is looking for –&lt;br /&gt;
someone who brings together that powerful mix of deep technical experience,&lt;br /&gt;
a strong personal financial situation, sound business judgment&lt;br /&gt;
and the ability to work productively with others.&lt;br /&gt;
But changing the world takes amazing people and I am prepared to find that person.&lt;br /&gt;
Please send me your resume at &#039;&#039;&#039;stephanie@sarka.us&#039;&#039;&#039; with a few words about your areas of expertise.&lt;br /&gt;
  &lt;br /&gt;
p.s.  No web shops or other third party providers, please.&lt;br /&gt;
&lt;br /&gt;
== Looking for a Partner/CTO ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 05-19-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
Equipped with a business idea in the semantic web world for enterprises and a business background,&lt;br /&gt;
I am looking for a partner/CTO to take the technical lead on the startup that I am working on.&lt;br /&gt;
This is an opportunity to get involved with a pre-funding start-up as a founding partner.&lt;br /&gt;
If you are available, open for a non-paid / equity job,&lt;br /&gt;
and are interested in getting involved in a start-up adventure, please get in touch.&lt;br /&gt;
I’m happy to discuss the project in greater detail.&lt;br /&gt;
&lt;br /&gt;
Please contact me at boaz_cn@hotmail.com (Boaz)&lt;br /&gt;
&lt;br /&gt;
== Chief Architect, Financial Times Search (FTS) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; April 13, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039; Chief Architect &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Reports To:&#039;&#039;&#039; Chief Technology Officer &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Stamford, CT &lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Here’s something to take note of if you are experienced in search and intelligent systems.   &lt;br /&gt;
&lt;br /&gt;
If you are creative, AND quantitative, we have an extraordinary opportunity for you.&lt;br /&gt;
Think of us as the place for which you got all of that education and experience.&lt;br /&gt;
If you think you might just be the right person who is interested in aiding people&lt;br /&gt;
with the burning questions of their day,&lt;br /&gt;
leveraged and fortified with sophisticated computer software&lt;br /&gt;
and have a demonstrated interest in AI or natural language processing,&lt;br /&gt;
we have a place for you to expand your horizons.   &lt;br /&gt;
&lt;br /&gt;
Financial Times Search (FTS)&lt;br /&gt;
(part of the Financial Times Group in turn part of Pearson the largest education publisher in the world)&lt;br /&gt;
is a new business utilizing a unique and proprietary search platform.&lt;br /&gt;
Our new stand-alone search product is called Newssift&lt;br /&gt;
and will soon be indexing tens of thousands of sources and many millions of articles for business people.&lt;br /&gt;
Our Beta is up and working at Newssift.com.   &lt;br /&gt;
 &lt;br /&gt;
This startup stands to deliver targeted search results&lt;br /&gt;
with a level of accuracy and relevance unmatched on the web today&lt;br /&gt;
and we would like to engage an experienced technical manager&lt;br /&gt;
to help plan and scope the engineering and significant product aspects of this business.&lt;br /&gt;
Check the product and the reviews out on the web and you will see we are on to something.    &lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Position Summary:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Responsible for architecture of a semantic based search product.&lt;br /&gt;
Researches, analyzes, and recommends technologies&lt;br /&gt;
to develop new and enhance existing systems to support semantic search applications.&lt;br /&gt;
Serves as a technical expert and is responsible for resolving complex problems&lt;br /&gt;
and guiding development to improve www.newssift.com, a consumer facing search application.&lt;br /&gt;
The role, while focusing on search,&lt;br /&gt;
will also interface with building systems around a subscription based product&lt;br /&gt;
as well as monitoring and measurement tools.&lt;br /&gt;
A command of building systems around where to go to answer a business person’s questions&lt;br /&gt;
is the main focus of the job.   &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Required Skills and Experience:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* Experience with mentoring software development resources.&lt;br /&gt;
* Experience with scoping the technical side of consumer facing products that require scale and contextual consumer experiences. &lt;br /&gt;
* Experience with understanding traffic flows and consumer choices.   Ideally the candidate has experience with mining historical consumer choice data to make a service more intelligent about consumer options. &lt;br /&gt;
* Experience in technical architectures for search applications, including data structures and algorithms to support entity extraction, disambiguation, normalization and clustering.&lt;br /&gt;
* Experience with architecture models and design patterns to support development and maintenance of taxonomies and knowledge bases.  Ideal candidate has been working on ways to do this in an open wiki way or by using data available from multiple sources. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Required Minimum Education:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* Master’s Degree in Computer Sciences or computational linguistics or equivalent field or related experience. &lt;br /&gt;
&lt;br /&gt;
Think about it.&lt;br /&gt;
One of the world’s largest companies, committed to education and information,&lt;br /&gt;
is on the leading edge of introducing a meaning based search and query tool.&lt;br /&gt;
Indeed we have a platform off of which to work, and a beachhead in the market.&lt;br /&gt;
If you have your sights set high – you should follow up with this one.   &lt;br /&gt;
&lt;br /&gt;
Only highly motivated and very smart folks need apply to Susan Blank at susankb48@msn.com&lt;br /&gt;
&lt;br /&gt;
==Web Tier/UI Developer (AJAX and JavaScript), Financial Times Search (FTS) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; April 13, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;          Web Tier/UI Developer (AJAX and JavaScript) &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Reports To:&#039;&#039;&#039;                 Senior Director Software Development &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039;                     Stamford, CT&lt;br /&gt;
&lt;br /&gt;
Financial Times Search (FTS) is a new business utilizing a unique and proprietary search engine&lt;br /&gt;
consistent with the heritage of the Financial Times.&lt;br /&gt;
FTS yields targeted search results with a level of accuracy and relevance unmatched on the web today.&lt;br /&gt;
FTS is a member of the Financial Times Group, part of Pearson PLC,&lt;br /&gt;
the world’s largest educational publisher&lt;br /&gt;
and owner of familiar businesses, including Prentice Hall and Penguin Books. &lt;br /&gt;
&lt;br /&gt;
Position Overview: &lt;br /&gt;
&lt;br /&gt;
We’re looking for a web tier/UI developer&lt;br /&gt;
with demonstrable experience developing high-performance dynamic user interfaces for consumer-facing web sites.&lt;br /&gt;
The ideal candidate will have the skills to develop, enhance and maintain a Rich Internet Application (RIA). &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Position Requirements:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* Expertise using JavaScript and AJAX to develop highly interactive, responsive user interfaces.&lt;br /&gt;
* Extensive knowledge of jQuery and JSON.&lt;br /&gt;
* Working knowledge of Java, JSP and XML.&lt;br /&gt;
* Excellent written and verbal communications skills.&lt;br /&gt;
* Experience in CSS.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Education Requirements:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* BS CS/EE or equivalent technical degree with a minimum of 3 years real-world experience.&lt;br /&gt;
&lt;br /&gt;
Think about it.&lt;br /&gt;
One of the world’s largest companies, committed to education and information,&lt;br /&gt;
is on the leading edge of introducing a meaning based search and query tool.&lt;br /&gt;
Indeed we have a platform off of which to work, and a beachhead in the market.&lt;br /&gt;
If you have your sights set high – you should follow up with this one.&lt;br /&gt;
&lt;br /&gt;
Only highly motivated and very smart folks need apply to Susan Blank at susankb48@msn.com&lt;br /&gt;
&lt;br /&gt;
==VP, Search and Semantic Technology. Elsevier Labs - New York==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; Mar 19, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
VP, Search and Semantic Technology&amp;lt;br&amp;gt;&lt;br /&gt;
Elsevier Labs &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Apply online to this position.&lt;br /&gt;
https://reedelsevier.taleo.net/careersection/51/jobdetail.ftl?lang=en&amp;amp;job=40524&lt;br /&gt;
&lt;br /&gt;
Elsevier is currently seeking a VP, Search and Semantic Technology&lt;br /&gt;
to identify with the needs of electronic publishing&lt;br /&gt;
and define a search and discovery technology strategy for next generation electronic products.&lt;br /&gt;
This position works closely with Elsevier product management, Labs, Enterprise Architecture,&lt;br /&gt;
and Elsevier divisional strategy&lt;br /&gt;
to ensure that the semantic technologies implement and inform the product vision.&lt;br /&gt;
&lt;br /&gt;
Responsibilities include:&lt;br /&gt;
&lt;br /&gt;
* Test, implement, program and evaluate search engines and partnerships&lt;br /&gt;
* Keep current with cutting edge research in semantic technologies and translate these into practical timelines for Elsevier&lt;br /&gt;
* Analyze implications of integrating new semantic technologies from both a technical perspective and a user/customer perspective.&lt;br /&gt;
* Work with product strategy to inform and help shape step changes in customer value&lt;br /&gt;
* Work with engineering development managers on assigned projects to plan integration of new technologies&lt;br /&gt;
* Define solutions and evaluate trade-offs for epublishing needs as related to search and other semantic capabilities&lt;br /&gt;
* Work with product architects to ensure that the search technologies meets the specific business needs of future products&lt;br /&gt;
* Manage, develop and own high level technical proposals and effort estimates as the initial piece of the overall product process&lt;br /&gt;
* Determine and recommend technical skills needed to complete a project&lt;br /&gt;
* Responsible for the knowledge transfer of information throughout Elsevier and with the engineering product groups regarding technical issues as related to search and semantic technologies&lt;br /&gt;
* Significantly contribute to knowledge base throughout Elsevier in the form of technical talks, white papers and seminars on technology that support the electronic publishing efforts throughout the business units&lt;br /&gt;
* Contribute to overall platform and technology directions for Elsevier with an emphasis on search&lt;br /&gt;
* Provide Consulting role to projects and product groups for the key search-related technology issues for their business&lt;br /&gt;
* Significantly contribute to RE Ventures&#039; evaluations of emerging search and discovery technologies&lt;br /&gt;
* Assist Enterprise Architects in coordination with RE Applied Technology for the planning and introduction of new search-related technologies&lt;br /&gt;
* Key communicator with IT and product management regarding current and future search capabilities&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;GENERAL KNOWLEDGE:&#039;&#039;&#039;&lt;br /&gt;
 &lt;br /&gt;
* Ability to influence&lt;br /&gt;
* Proven ability to communicate (written and verbal) technology effectively to senior management, product marketing and to engineers&lt;br /&gt;
* Proven ability to absorb large amounts of technical and business detail and synthesize that into a usable problem definition and technical approach&lt;br /&gt;
* Ability to work collaboratively, by directing and guiding the technical direction of a project&lt;br /&gt;
* Ability to work with Senior level development managers&lt;br /&gt;
* Ability to work with ambiguous situations and bring them to closure&lt;br /&gt;
* Ability to influence without direct management&lt;br /&gt;
* Exceptional written and oral communication and presentation skills&lt;br /&gt;
* Demonstrated leadership skills&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;TECHNICAL SKILLS:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Practical and theoretical experience about information retrieval systems&lt;br /&gt;
* Specific technical expertise in these areas:&lt;br /&gt;
* Concept classification, XML, vector space searching, probabilistic retrieval and neural networks&lt;br /&gt;
* Programming techniques: parsing, syntactic analysis, semantic analysis, use of thesauri and ontologies&lt;br /&gt;
* Query processing and query understanding including question/answer paradigms, question extraction, multiple-constraint search paradigms&lt;br /&gt;
* Ability to work with large and diverse data sets&lt;br /&gt;
* Interoperability of search across multiple engines and sites (federated or meta-search) including query normalization&lt;br /&gt;
* PhD in Information Retrieval highly desirable&lt;br /&gt;
* Masters degree in computer science and/or 7+ years of relevant experience in engineering; 5+ years in internet/web technologies&lt;br /&gt;
* Experience as architect (or design lead) in a significant project&lt;br /&gt;
* Experience in transitioning research and prototypes into production a plus&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Other Locations:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* United States-Pennsylvania-Philadelphia&lt;br /&gt;
* United States-Maryland-Rockville&lt;br /&gt;
* United States-California-Irvine&lt;br /&gt;
* United States-Missouri-St Louis&lt;br /&gt;
* United States-New Jersey-Bridgewater&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Closing Date:&#039;&#039;&#039; Ongoing&lt;br /&gt;
&lt;br /&gt;
https://reedelsevier.taleo.net/careersection/51/jobdetail.ftl?lang=en&amp;amp;job=40524&lt;br /&gt;
&lt;br /&gt;
== Funded Startup is Looking for a Experienced Java Engineer ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 02-25-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
AdaptiveBlue (http://getglue.com) is an innovative, well funded semantic web startup in New York City,&lt;br /&gt;
named among the 250 best startups around the world by AlwaysOn Network.&lt;br /&gt;
AdaptiveBlue is focused on results in a fast-paced environment.&lt;br /&gt;
&lt;br /&gt;
We are looking for talented, smart, hard working, experienced and passionate Java Engineer&lt;br /&gt;
to help us build the next generation of web browsing technologies.&lt;br /&gt;
&lt;br /&gt;
You will be working on AdaptiveBlue&#039;s back end - web service, database, and metrics.&lt;br /&gt;
This is an exciting opportunity for someone interested in the semantic technologies,&lt;br /&gt;
skilled with algorithms, and proficient in Java.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Must love coding and hard work&lt;br /&gt;
* B.S. or higher in CS, Math or engineering&lt;br /&gt;
* At lest 5 years of experience as a software engineer&lt;br /&gt;
* At least 5 years of experience in Java programming&lt;br /&gt;
* At least 3 years of experience with SQL&lt;br /&gt;
* Strong knowledge of basic data structures and algorithms&lt;br /&gt;
* Strong knowledge of design patterns, refactoring and unit testing&lt;br /&gt;
* Strong knowledge of concurrent programming&lt;br /&gt;
* Understanding of distributed, large-scale systems&lt;br /&gt;
* Experience with XML/XSL and REST-based Web Services&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The following are a plus, but not required:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Experience with Amazon Web Services&lt;br /&gt;
* Knowledge of semantic markups and ontologies&lt;br /&gt;
* Experience with semantic APIs like Calais&lt;br /&gt;
* Experience with SQL query tuning&lt;br /&gt;
&lt;br /&gt;
We offer competitive salary, full benefits, 401k, Awesome Mac hardware,&lt;br /&gt;
Herman Miller chairs, and Fresh Direct snacks.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;If you love coding and want to work on exciting things&lt;br /&gt;
that change the way people interact with the web, please send us all of the items below:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* A cover letter describing why you are a good fit&lt;br /&gt;
* A resume with your experiences&lt;br /&gt;
* A sample of Java code that you have written in the past year&lt;br /&gt;
&lt;br /&gt;
* Compensation: Solid base + 401k + stock options&lt;br /&gt;
* Principals only. Recruiters, please don&#039;t contact this job poster.&lt;br /&gt;
* Please, no phone calls about this job!&lt;br /&gt;
* Please do not contact job poster about other services, products or commercial interests.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; job-1049787041@craigslist.org&lt;br /&gt;
&lt;br /&gt;
== Full-time Permanent Semantic Java Software Engineer ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 02-24-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Category:&#039;&#039;&#039; &lt;br /&gt;
Semantic Java Software Engineer&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; &lt;br /&gt;
Boston, MA,  USA - Metro/West&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039; &lt;br /&gt;
We are seeking a full-time permanent semantic java software engineer for our client in the Metro-Boston area.&lt;br /&gt;
If you are interested in building a semantic engine -&lt;br /&gt;
do you consider yourself an ontology expert?&lt;br /&gt;
Know RDF and/or OWL inside and out?&lt;br /&gt;
We&#039;ve got a cool project converting a .Net platform into a Java platform.&lt;br /&gt;
We need a strong coder...&lt;br /&gt;
someone who really really enjoys coding with commercial products experience and a great attitude.&lt;br /&gt;
This you? Give us a call.&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Salary:&#039;&#039;&#039; Open depending on experience&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; &lt;br /&gt;
Sarah Bevelaqua|Technical Recruiter&amp;lt;br&amp;gt;&lt;br /&gt;
The FootBridge Companies |Direct Placement Services Group&amp;lt;br&amp;gt;&lt;br /&gt;
http://www.FootBridgeDirect.com&amp;lt;br&amp;gt;&lt;br /&gt;
Office: 978.474.4455| Toll Free: 877.807.8400&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Artificial Intelligence Software Developer, Washington DC Area==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 02-10-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Category:&#039;&#039;&#039; Artificial Intelligence Software Developer&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Washington, DC, USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Clearance:&#039;&#039;&#039; US DoD Secret required.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Experience:&#039;&#039;&#039; Five years experience in Artificial Intelligence related programming&lt;br /&gt;
or a Master&#039;s Degree in Artificial Intelligence or a related field. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills:&#039;&#039;&#039; Experience in the following, or similar, subjects a plus:&lt;br /&gt;
Software Patterns, Agent-Based Simulation, Software Optimization and Scalability,&lt;br /&gt;
Open Source Contributions, Ontologies, Inference Engines, Evolutionary Computation,&lt;br /&gt;
Neural Networks, Bayesian Networks, Fuzzy Expert Systems, Data Mining,&lt;br /&gt;
Case-Based Reasoning, Game Trees and Game Theory, Social Network Analysis,&lt;br /&gt;
Statistical Design of Experiments. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Programming Language:&#039;&#039;&#039; Java&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Other Software and Development Tools:&#039;&#039;&#039;&lt;br /&gt;
Experience in the following, or similar, software a plus -&lt;br /&gt;
Protégé, Owl, Pellet, Jena, Jastor, Weka, Repast, Groovy, ECJ, Ptolemy,&lt;br /&gt;
JFuzzyLogic, Joone, Jung, R, Jboss, Prefuse&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;&lt;br /&gt;
Artificial Intelligence Software Developers needed&lt;br /&gt;
for cutting-edge Computational Social Science simulation project.&lt;br /&gt;
Will be working in a Team-based environment with other SW developers.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Potential Task areas include:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Scientific Experimentation and Data Mining:  Enhance existing software to support finding patterns in wargame events, for the purposes of testing hypotheses about relations between events and for discovering relationships between events.   Incorporate open source data mining and artificial intelligence software, such as, for example, Weka and ECJ, to automatically find patterns in moves and outcomes&lt;br /&gt;
&lt;br /&gt;
* Interoperation of Hybrid Models:  Support the interoperation of hybrid models, to ensure that models that have multiple resolutions and perspectives share meaning.   Through xml, implement the translations between models.  Integrate with integration tools that support semantic interoperation through ontologies, and other methodologies such as, for example, the COMPOEX backplane or Ptolemy. Help implement a Hub and Spoke design for translation between data models a system by which simulation models and data of different data models may interoperate through their own individual data models, a common data model, and a translation data model between their own individual data models and the common data model. Enhance the software enforcement of a data model &amp;quot;contract&amp;quot; using ontology-based software engineering techniques.  Extend an existing example implemented in Jena and Jastor, converting it to a Dynamic Object Model (DOM) language (such as, for example, Groovy).&lt;br /&gt;
&lt;br /&gt;
* Support for Conflict Resolution between models.  Support the conflict resolution of possibly conflicting hybrid models, for the purposes of making a unified, coherent picture of the social environment.&lt;br /&gt;
Integrate open source software (such as, for example, JFuzzyLogic) to match the simulation output to correlative social study data, and establish the correlative relations that should exist between and within component social science models in support of validation and consensus building of possibly conflicting models.&lt;br /&gt;
Implement a framework for consensus-building.  Make possible the specification of arbitrary schemes for developing a model consensus.  The framework would have a way to handle issues of integration, for example, the models may conflict by having mutually exclusive or uncorrelated behaviors.   The framework would allow various consensus schemes to be switched in and out.  Implement a simple prototype weighted voting scheme in the framework&lt;br /&gt;
&lt;br /&gt;
* Support for Component Models.  Enhance the Nexus Intelligent Adaptive Agent Based Models to increase their generality, scalability, efficiency, and ability to work as component models for other software.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; &amp;lt;br&amp;gt;&lt;br /&gt;
Debbie Duong &amp;lt;br&amp;gt;&lt;br /&gt;
debbieduong62  at gmail.com&lt;br /&gt;
&lt;br /&gt;
==Knight Professor of the Practice of Journalism and Public Policy Studies, Duke University==&lt;br /&gt;
&lt;br /&gt;
Duke University&amp;lt;br&amp;gt;&lt;br /&gt;
Sanford Institute of Public Policy&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Durham, NC, USA&lt;br /&gt;
&lt;br /&gt;
The Sanford Institute of Public Policy at Duke University seeks applicants&lt;br /&gt;
for the Knight Professor of the Practice of Journalism and Public Policy Studies,&lt;br /&gt;
an endowed chair in Duke’s DeWitt Wallace Center for Media and Democracy.&lt;br /&gt;
&lt;br /&gt;
We seek a person who will help in the development of a new field, computational journalism.&lt;br /&gt;
Advances in data availability, technology, and algorithms offer the prospect&lt;br /&gt;
that part of the media’s watchdog function may be supplemented by analyses done by computers.&lt;br /&gt;
The Knight Professor at Duke will help translate advances&lt;br /&gt;
in areas such as artificial intelligence and the semantic web&lt;br /&gt;
into the development of products that help journalists and other community members&lt;br /&gt;
hold institutions accountable.&lt;br /&gt;
In part, this may involve computerizing aspects of investigative reporting.&lt;br /&gt;
&lt;br /&gt;
Candidates for the chair should have experience in working&lt;br /&gt;
at the intersection of technology and information generation.&lt;br /&gt;
They may be working in computer assisted reporting, or visualization of data on the Web,&lt;br /&gt;
or the development of algorithms to sift through publicly available data&lt;br /&gt;
for clues to the performance of government.&lt;br /&gt;
Ideal candidates could include reporters or editors involved in innovations in digital news ventures&lt;br /&gt;
or programmers involved in the development of artificial intelligence or semantic web products&lt;br /&gt;
aimed at news gathering and reporting.&lt;br /&gt;
Candidates should be strongly committed to advancing the development of computational journalism as a field&lt;br /&gt;
through innovative research and the development of new digital tools.&lt;br /&gt;
&lt;br /&gt;
The Knight Professor will teach courses in computational journalism and public policy&lt;br /&gt;
in our undergraduate public policy major and certificate program in policy journalism and media studies.&lt;br /&gt;
The Knight Professor will play a key role in the teaching, research and policy engagement activities&lt;br /&gt;
of the DeWitt Wallace Center for Media and Democracy.&lt;br /&gt;
&lt;br /&gt;
Candidates for this position should send a CV and other materials to:&lt;br /&gt;
&lt;br /&gt;
Professor James T. Hamilton&amp;lt;br&amp;gt;&lt;br /&gt;
Knight Search Committee Chair&amp;lt;br&amp;gt;&lt;br /&gt;
DeWitt Wallace Center for Media and Democracy&amp;lt;br&amp;gt;&lt;br /&gt;
Sanford Institute of Public Policy&amp;lt;br&amp;gt;&lt;br /&gt;
Duke University&amp;lt;br&amp;gt;&lt;br /&gt;
Box 90241&amp;lt;br&amp;gt;&lt;br /&gt;
Durham, NC 27708-0241&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Review of applications will begin immediately.&lt;br /&gt;
Applications received by December 31, 2008 will be guaranteed full consideration.&lt;br /&gt;
&lt;br /&gt;
Duke University is an Equal Opportunity/Affirmative Action Employer.&lt;br /&gt;
&lt;br /&gt;
==Executive Director Emerging Platforms ORGANIC, INC.==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 10-23-08&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Area code:&#039;&#039;&#039; 212&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Pay rate:&#039;&#039;&#039; open&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Experience:&#039;&#039;&#039; Exceptional Experience.&lt;br /&gt;
&lt;br /&gt;
Organic is all about the exceptional.&lt;br /&gt;
We are a leading digital communications agency –&lt;br /&gt;
the first, in fact – focused on designing and building exceptional experiences&lt;br /&gt;
that help make the online channel really work for leading companies.&lt;br /&gt;
We work on everything digital – from websites to online marketing campaigns,&lt;br /&gt;
from mobile applications to digital billboards –&lt;br /&gt;
for companies such as Chrysler LLC, Warner Bros. International, Geek Squad and Bank of America.&lt;br /&gt;
We are passionate about emerging platforms,&lt;br /&gt;
and this helps keep us on the edge of the hottest and most interesting stories in the industry.&lt;br /&gt;
And, as a leader in the marketplace, we have the opportunity to work on really exciting projects. &lt;br /&gt;
&lt;br /&gt;
Organic has an open, diverse culture that encourages participation and innovation.&lt;br /&gt;
We reward exceptional work and creative ideas.&lt;br /&gt;
At Organic we believe in working hard and having fun while we’re doing it.&lt;br /&gt;
We have great benefits and perks including massages, pet insurance, Wednesday bagels and Friday cocktails. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* The Executive Director of Emerging Platforms leads client innovation in the areas of new media technologies, platforms and networks. &lt;br /&gt;
* Oversees the Emerging Platform team which is comprised of strategist and interactive technology developers. &lt;br /&gt;
* Provides clients with an outward looking perspective of the new media landscape and demonstrates what clients should expect and how these developments may impact their online experience, business, industry, competition and customers. &lt;br /&gt;
* Leads the rapid prototyping initiative for Organic. &lt;br /&gt;
* Develop internal projects and employee education programs to demonstrate new technologies to Organic employees with the hope of eventual client education and suggestion. &lt;br /&gt;
* Builds and manages connections to the new media industry to identify new opportunities for Organic clients and Organic employees (toolsets, frameworks and development technologies). &lt;br /&gt;
* Industry perspective and thought leadership platforms for Organic and for industry media (corporate marketing). &lt;br /&gt;
* Attendance and Organic&#039;s representation at industry conferences and events. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Education and Work Experience:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* 10 yrs overall experience preferred including: &lt;br /&gt;
* 10+ yrs experience working in marketing strategy, account planning and interactive marketing. &lt;br /&gt;
* Experience in building, growing and leading successful teams (either in an office, practice, or global teams of 50+). &lt;br /&gt;
* Client-focused – builds long term high-level relationships that reflect well on Organic, proven history in business management and growth. &lt;br /&gt;
* Ability to manage and be held accountable for resources, revenues and budgets. &lt;br /&gt;
* Understands and can both, manage and work within a matrix management organization. &lt;br /&gt;
* Personable; very poised; a strong conceptual thinker; have high energy; exhibit excellent writing, communication and presentation skills. &lt;br /&gt;
* Inspirational leader – exhibits integrity, and respectful interactions with past recognizable clients, superiors, peers and subordinates. &lt;br /&gt;
* Proven history of growing a practice or major service. &lt;br /&gt;
* Able to thrive in a &amp;quot;non-traditional&amp;quot;, entrepreneurial environment. &lt;br /&gt;
* Availability and willingness to travel 35%of time. &lt;br /&gt;
* Bachelors/MBA degree preferred or equivalent work experience. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Travel required:&#039;&#039;&#039; 35%&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Telecommute:&#039;&#039;&#039; no&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; Ross, http://www.organic.com, ross @ organic . com&lt;br /&gt;
&lt;br /&gt;
==Web 2.0 Build Out==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 10-22-2008&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
What if there was absolutely nothing stopping you from inventing the next big thing?&lt;br /&gt;
What would you do if you had a clear voice in developing the next generation of web based information services?&lt;br /&gt;
How would you respond if you had the opportunity to be a creative force&lt;br /&gt;
in one of the most unique and inspiring work cultures in NYC? &lt;br /&gt;
&lt;br /&gt;
Our client is a startup web unit within a global telecom leader,&lt;br /&gt;
launching a mobile and web based platform&lt;br /&gt;
that will revolutionize the way consumers interact with and manage information.&lt;br /&gt;
They are early in development and building a team of innovative, intelligent professionals who love what they do.&lt;br /&gt;
We are looking for people who are obsessed with web technology, creative in their process&lt;br /&gt;
and can thrive in an entrepreneurial, fast paced environment.&lt;br /&gt;
Our ideal candidates work very well within a team, are self-motivated,&lt;br /&gt;
have a high level of energy and a strong drive to succeed.&lt;br /&gt;
They must have a passion for the internet and be current with cutting edge web trends and technologies.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Primary Responsibilities:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Work as key member of a small team developing foundational systems and services for an outstanding web 2.0 platform &lt;br /&gt;
* Work with vendors and agencies to integrate products and services into platform &lt;br /&gt;
* Hands on development of web systems and projects &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Development experience in standard internet technologies &lt;br /&gt;
* Experience in an SOA environment is preferred &lt;br /&gt;
* Linux deployment &lt;br /&gt;
* Solid Python development skills &lt;br /&gt;
* System design experience preferred &lt;br /&gt;
* Track record of successful releases &lt;br /&gt;
* Experience with large scale distributed systems &lt;br /&gt;
* RESTful services a plus &lt;br /&gt;
* JavaScripting, JSON, XML, Ruby, Erlang all a plus&lt;br /&gt;
&lt;br /&gt;
Interested candidates email resumes to jim@execuseek.net&lt;br /&gt;
&lt;br /&gt;
==Program Manager with Ontology and Taxonomy/OWL==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills:&#039;&#039;&#039; Taxonomy Ontology tools OWL+ Media Endeca&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 7-1-2008&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Area code:&#039;&#039;&#039; 212&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Pay rate:&#039;&#039;&#039; open&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
This individual will be responsible for multiple aspects of simultaneous long term projects including:&lt;br /&gt;
Consolidating project plans, responsibility for deadlines, ongoing support,&lt;br /&gt;
implementation and coordination and overall delivery of multiple inter-related projects.&lt;br /&gt;
&lt;br /&gt;
Accountable for the on time and on budget completion of these projects,&lt;br /&gt;
the Program Manager will define, coordinate and lead a blended team&lt;br /&gt;
thus requiring regular interaction with business, editorial, production, marketing, design and development factions&lt;br /&gt;
of the internet organizations.&lt;br /&gt;
The individual is expected to regularly participate in all phases of project development&lt;br /&gt;
with a strong team-oriented attitude.&lt;br /&gt;
Strong communication skills, both written and verbal are essential,&lt;br /&gt;
as is the ability to simultaneously manage multiple projects in a dynamic, challenging, and fast-paced environment.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Candidates will possess a Bachelor&#039;s Degree in a related field&lt;br /&gt;
* 5-7 years Project Management Experience with solid experience working in a cross-functional environment.&lt;br /&gt;
* This person will be a **manager of project managers** and should have experience pulling together very large programs of projects with multiple large interrelated subprojects&lt;br /&gt;
* Experience with search technologies is a must; specific experience with Endeca a major plus&lt;br /&gt;
* Experience with taxonomy / ontology tools a must; specific experience with OWL-based technologies a major plus&lt;br /&gt;
* Requires a high-level understanding of the integration of search and taxonomy/ontology tools with a content management system&lt;br /&gt;
* Online Publishing / Media experience a huge plus&lt;br /&gt;
* GREAT interpersonal skills&lt;br /&gt;
** this person will interface with the project managers, functional managers, and will be required to have frequent written and face-to-face communications with senior management&lt;br /&gt;
** Managing software vendor relationships and coordinates training, documentation, and communication with IT and the Titles.&lt;br /&gt;
* Creates and executes project work plans and revises as appropriate to meet changing needs and requirements.&lt;br /&gt;
* Creates and executes an aggregate program view of interrelated projects and has the ability to recognize, surface, and resolve issues across the program.&lt;br /&gt;
* PMI Certification is preferred, but not mandatory.&lt;br /&gt;
* Experienced with Web based projects, preferably in the online content/publishing space.&lt;br /&gt;
* Effectively applies our methodology and enforces project standards with an eye on emerging industry practices.&lt;br /&gt;
* Prepares for engagement reviews and quality assurance procedures.&lt;br /&gt;
* Recognizes and minimizes exposures and risks on the individual projects and across the program of projects.&lt;br /&gt;
* Ensures project documents are complete, current, and stored appropriately.&lt;br /&gt;
* Facilitates team and client meetings effectively, documents meeting minutes and agreements.&lt;br /&gt;
* Effectively communicates relevant project information to superiors in an engaging, informative, well-organized format. Must be able to effectively tailor communications of complex issues to audiences with varying levels of subject matter expertise.&lt;br /&gt;
* Acquires / possesses a thorough understanding of our capabilities.&lt;br /&gt;
* Resolves and/or escalates issues in a timely fashion and understands how to communicate difficult/sensitive information tactfully.&lt;br /&gt;
* Motivates team to work together in the most efficient manner.&lt;br /&gt;
* Suggests areas for improvement in internal processes along with possible solutions.&lt;br /&gt;
* Expert in MS Project, MS Excel, MS Powerpoint&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Travel required:&#039;&#039;&#039; none&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Telecommute:&#039;&#039;&#039; no&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Debbie Levy, [http://www.logiccorporation.com Logic]&lt;br /&gt;
&amp;lt;br&amp;gt;debbie @ logiccorporation . com&lt;br /&gt;
&lt;br /&gt;
== Semantic Web Senior Java Developer -- Alitora Systems ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
Alitora Systems is a start-up software company providing an innovative Semantic Search and Collaboration service.&lt;br /&gt;
&lt;br /&gt;
See: [http://www.alitora.com http://www.alitora.com]&lt;br /&gt;
&lt;br /&gt;
We are building the first true commercial Semantic Web Service as a SaaS platform.&lt;br /&gt;
We are primarily serving the Life Science Industry, such as the pharmaceutical industry,&lt;br /&gt;
which requires a huge amount of critical high value data&lt;br /&gt;
to support Drug Development, Competitive Intelligence, and Business Development.&lt;br /&gt;
&lt;br /&gt;
Our core technology is the kHarmony Semantic Database:&lt;br /&gt;
a general hyper-graph database used to store knowledge modeled semantically.&lt;br /&gt;
&lt;br /&gt;
We are seeking a senior developer with significant Java experience, including:&lt;br /&gt;
&lt;br /&gt;
* Strong Experience with Semantic technologies and/or Search Engine technologies, such as:&lt;br /&gt;
**Java Lucene search engine&lt;br /&gt;
**Machine Learning algorithms and toolkits&lt;br /&gt;
**JENA and related Java toolkits&lt;br /&gt;
&lt;br /&gt;
* Significant Experience implementing efficient algorithms in areas such as:&lt;br /&gt;
** Search engine query processing and/or index construction&lt;br /&gt;
** Inference algorithms, Logic Algorithms&lt;br /&gt;
** Machine Learning / Statistical Methods&lt;br /&gt;
** Graph Theoretic Algorithms&lt;br /&gt;
** Natural Language Processing algorithms, such as Entity Extraction&lt;br /&gt;
&lt;br /&gt;
* Understanding of Semantic Web concepts, such as:&lt;br /&gt;
** RDF, OWL, Inference Engines, URIs&lt;br /&gt;
&lt;br /&gt;
* Understanding of Graph Theory concepts, such as:&lt;br /&gt;
** Nodes, Edges, Depth-First-Search, Cliques&lt;br /&gt;
&lt;br /&gt;
* Appreciation of the Life Sciences&lt;br /&gt;
&lt;br /&gt;
* Great inter-personal / teamwork skills and communication skills&lt;br /&gt;
&lt;br /&gt;
* Able to prosper in a cutting-edge start-up environment&lt;br /&gt;
&lt;br /&gt;
* And especially, Passionate Enthusiasm for New Technology and Start-Ups&lt;br /&gt;
&lt;br /&gt;
---------------&lt;br /&gt;
&lt;br /&gt;
Please reply with a resume.&lt;br /&gt;
&lt;br /&gt;
We’re seeking to bring developers on as consultants or employees, depending on individual circumstance&lt;br /&gt;
&lt;br /&gt;
We&#039;re based in New York City, although some developers work remotely. We have daily meetings EST.&lt;br /&gt;
&lt;br /&gt;
Email: &#039;&#039;marc@alitora.com&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
==Experienced Java Programmer for Semantic R&amp;amp;D Position (Washington, DC)==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date: ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Washington, DC, USA&lt;br /&gt;
&lt;br /&gt;
Small, established startup (http://clarkparsia.com/) in downtown DC&lt;br /&gt;
seeks an experienced, systems-level Java programmer&lt;br /&gt;
to work on R&amp;amp;D projects and production systems in semantic technologies,&lt;br /&gt;
including reasoning, planning, description logics, semantic web services, logistics, etc.&lt;br /&gt;
Join the team responsible for Pellet, the leading OWL DL reasoner.&lt;br /&gt;
&lt;br /&gt;
The successful applicant will be smart and have a demonstrable record of getting things done.&lt;br /&gt;
&lt;br /&gt;
* Bachelor&#039;s or Master&#039;s in CS or relevant field&lt;br /&gt;
* Minimum 7 years of serious programming experience&lt;br /&gt;
* Experience with AI, KR, logic programming, or planning a definite plus&lt;br /&gt;
* Enthusiasm and ability for solving difficult, algorithmically novel problems required&lt;br /&gt;
* Familiarity with database theory valuable&lt;br /&gt;
* Familiarity with systems or operations research also valuable&lt;br /&gt;
* Excellent written and verbal communication skills&lt;br /&gt;
* Must be eligible to work permanently in the US&lt;br /&gt;
&lt;br /&gt;
We offer the following benefits:&lt;br /&gt;
&lt;br /&gt;
* Competitive salary and benefits (401k, FSA, major med, dental, vision, etc)&lt;br /&gt;
* A chance to build new, cool stuff that people use&lt;br /&gt;
* Interesting, vibrant work environment: &lt;br /&gt;
** Free lunch, &lt;br /&gt;
** baseball and other sports outings, &lt;br /&gt;
** Wii tournaments,&lt;br /&gt;
** great espresso &amp;amp; coffee &lt;br /&gt;
** Beer Fridays&lt;br /&gt;
** 1 block from Metro (Mt Vernon Square, Green Line) near Convention Center&lt;br /&gt;
&lt;br /&gt;
To apply, send resume and a cover letter to Kendall Clark, kendall@clarkparsia.com.&lt;br /&gt;
&lt;br /&gt;
==Ontology Modeler Systems Engineer, Mountain View CA==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Mountain View, CA, USA&lt;br /&gt;
&lt;br /&gt;
My client is in Mountain View CA.&lt;br /&gt;
They are looking for an Ontology Modeler/Systems Engineer.&lt;br /&gt;
The primary focus will be on developing information models for Aerospace Vehicles and Systems.&lt;br /&gt;
This includes general enterprise architecture modeling (e.g., organizations, processes, tools)&lt;br /&gt;
as well as models specific to engineering domains (e.g., vehicles, sub-systems, devices and functions).&lt;br /&gt;
You will be using Semantic Web standards (RDF and OWL) and XML.&lt;br /&gt;
Any knowledge of these technologies is a big plus, but we will train the right person.  &lt;br /&gt;
&lt;br /&gt;
This position requires:&lt;br /&gt;
&lt;br /&gt;
* More than 5 years of modeling experience (either object modeling, system modeling, data modeling or knowledge modeling)&lt;br /&gt;
* A degree and/or work experience in an engineering field&lt;br /&gt;
* Strong communications skills (experience interviewing people for knowledge capture, running workshops, etc.)&lt;br /&gt;
&lt;br /&gt;
Knowledge of space systems and their engineering disciplines is a distinct advantage.&lt;br /&gt;
This includes avionics, mechanics, hydraulics, propulsion, guidance and navigation,&lt;br /&gt;
telemetry and control systems. Some past or current programming skills will be an advantage,&lt;br /&gt;
especially in Java, Prolog, or a functional language such as Haskell.&lt;br /&gt;
&lt;br /&gt;
The ideal candidate would also have one (or more) of the following qualifications.&lt;br /&gt;
Knowledge and experience of:&lt;br /&gt;
&lt;br /&gt;
* Modeling formalisms such as UML and SysML.&lt;br /&gt;
* Information and knowledge structuring formalisms such as ASN.1, XML, RDF, OWL.&lt;br /&gt;
* Enterprise Architecture frameworks such as DODAF and TOGAF&lt;br /&gt;
* Training and/or experience in computational  linguistics&lt;br /&gt;
&lt;br /&gt;
The person should enjoy working on challenging problems, be a self starter,&lt;br /&gt;
have strong communication skills and be ready to show a lot of initiative.&lt;br /&gt;
He/she should want to work in a dynamic and growing small company&lt;br /&gt;
with collegial culture and many opportunities to learn and do different things. &lt;br /&gt;
&lt;br /&gt;
We offer competitive salary, bonus, major medical, dental, and an attractive stock option plan.&lt;br /&gt;
The position is in our Mountain View, California location.&lt;br /&gt;
We will provide assistance with the relocation expenses.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Mary Frances Hunter&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
503 232 8822&amp;lt;br&amp;gt;&lt;br /&gt;
&#039;&#039;Recruiter Extraordinaire&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Software Engineer, Mountain View CA==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Mountain View, CA, USA&lt;br /&gt;
&lt;br /&gt;
My client is in Mountain View CA. They are looking for Software Engineers who know server-side Java development and want to work on a Semantic Web product.&lt;br /&gt;
&lt;br /&gt;
The Product Suite is built on Eclipse but with a Web UI. They want to add more features. This involves design, development, test and integration of the new features. We need someone who is a self-starter, as there is little supervision. There is collaboration.&lt;br /&gt;
&lt;br /&gt;
New feature may include the importing and exporting of data, merging and transferring data, connecting to external data bases, managing changes to forms…..&lt;br /&gt;
&lt;br /&gt;
Technologies involved are: Java, Eclipse, Topcat, Adobe Products, AJAX/Flex, Mash-ups, OWL, RDF/S, SPARQL… &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Mary Frances Hunter&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
503 232 8822&amp;lt;br&amp;gt;&lt;br /&gt;
&#039;&#039;Recruiter Extraordinaire&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Senior Engineer, Semantic Web Datastore ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Cambridge, MA, USA&lt;br /&gt;
&lt;br /&gt;
ITA Software&lt;br /&gt;
&lt;br /&gt;
http://www.itasoftware.com/careers/jlisting.html?jid=26&lt;br /&gt;
&lt;br /&gt;
ITA Software -- known for its algorithm-intensive airfare search and reservations products --&lt;br /&gt;
has also been doing novel Semantic Web and Data Integration work since 2004.&lt;br /&gt;
We&#039;re now turning the corner from research to deployment&lt;br /&gt;
and seek the right person to engineer our data store for scalability and performance.&lt;br /&gt;
Follow the link above for details.&lt;br /&gt;
&lt;br /&gt;
[Thanks, Marco, for inviting me to post here!  -JustinITA]&lt;br /&gt;
&lt;br /&gt;
=Archive=&lt;br /&gt;
&lt;br /&gt;
==Information Architect, Collection Information &amp;amp; Access, J. Paul Getty Museum==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; December 2008&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Los Angeles, CA, USA&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The department of Collection Information &amp;amp; Access at the J. Paul Getty Museum&lt;br /&gt;
is seeking an Information Architect to oversee the back-end structure, data models, systems and applications&lt;br /&gt;
used by the Museum to support the management and dissemination of documentation, digital assets, and metadata&lt;br /&gt;
on the collection and to ensure its accessibility in the networked environment.&lt;br /&gt;
The Information Architect will lead efforts in restructuring the way information&lt;br /&gt;
is stored, systems integrated, and data published&lt;br /&gt;
so as best to ensure efficiency in processes, scalability and sustainability, and resource discovery.&lt;br /&gt;
This will involve architectural designs, analysis, integration, and strategic direction&lt;br /&gt;
for how best to manage existing enterprise-wide applications&lt;br /&gt;
such as collections management, content management and digital asset management systems&lt;br /&gt;
with other custom grown applications and open source solutions,&lt;br /&gt;
in addition to overseeing data modeling and strategies that are system independent.&lt;br /&gt;
The position will be responsible for the maintenance of data models, data dictionaries, and processes;&lt;br /&gt;
work with technical staff across the Getty to build mechanisms&lt;br /&gt;
for exchanging data and metadata between repositories;&lt;br /&gt;
and work closely with user communities&lt;br /&gt;
for requirements analysis, problem definition and solutions development. &lt;br /&gt;
&lt;br /&gt;
The ideal candidate will utilize standards, best practices, and forward-thinking solutions&lt;br /&gt;
for structuring the Museum&#039;s information architecture,&lt;br /&gt;
and be able to provide analysis, documentation, and ROI for strategies.&lt;br /&gt;
The candidate should have experience in all phases of the software development cycle;&lt;br /&gt;
understand and be technically proficient in the environments&lt;br /&gt;
in which software applications operate (i.e. Unix, Windows);&lt;br /&gt;
have familiarity with semantic technologies&lt;br /&gt;
including triple stores, natural language processing, and clustering techniques.&lt;br /&gt;
The candidate should be comfortable with writing technical documentation and design documents,&lt;br /&gt;
outlining detailed process flow and workflow mappings, have strong analytical skills,&lt;br /&gt;
excellent oral and written communication skills,&lt;br /&gt;
and the ability to effectively work in a team environment. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039; Proven experience working with relational databases (Oracle 10g), SQL Server and using Structured Query Language; familiarity with &amp;quot;C++&amp;quot;, JAVA or similar object-oriented programming language; and proficient at UNIX scripting languages; JavaScript, HTML, CSS, XML and XSLT. Working knowledge of ontologies and ontology standards like RDF and concepts associated with the Semantic Web.  &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications:&#039;&#039;&#039; Bachelor&#039;s Degree in Computer Science, Library &amp;amp; Information Science, Information Technology, or related studies required, Master&#039;s preferred. Minimum 8 years of experience in the electronic management of information, and developing, implementing and managing information architecture in a publishing, library, or educational repository environment strongly preferred. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Please email cover letter and resume to jobs@getty.edu indicating in the subject line, &amp;quot;Museum Information Architect /AT.&amp;quot; OR send to: The J. Paul Getty Trust, 1200 Getty Center Drive, Suite 400, Los Angeles, Ca 90049-1681 and reference &amp;quot; Museum Information Architect /AT&amp;quot; in your cover letter. No phone calls, please. EOE.&lt;br /&gt;
&lt;br /&gt;
http://www.getty.edu/about/opportunities/tech_opps.html&lt;br /&gt;
&lt;br /&gt;
== Global Director of Semantic Technology Solutions ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Providing thought-leadership in the adoption of emerging semantic technologies as a means to add value to information products and thereby drive revenue.&lt;br /&gt;
* Overseeing software development of Synaptica® from Dow Jones, an enterprise-class taxonomy and ontology management software product.&lt;br /&gt;
* Leading the development of new taxonomies and metadata that can be used to enrich information content.&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Full Job Description online at:&lt;br /&gt;
&lt;br /&gt;
http://careers.peopleclick.com/jobposts/Client40_DowJones/BU1/External/pck314-6886.htm&lt;br /&gt;
&lt;br /&gt;
Dave Clarke&amp;lt;br&amp;gt;&lt;br /&gt;
Global Taxonomy Director&amp;lt;br&amp;gt;&lt;br /&gt;
Dow Jones&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Global Taxonomy Director ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Coordinate initiatives with the regional Client Solutions Directors to achieve and exceed Dow Jones Taxonomy Services revenue and operational goals, leveraging Dow Jones&#039; taxonomy and editorial expertise.&lt;br /&gt;
* Define and communicate Taxonomy capabilities, services and pricing models.&lt;br /&gt;
* Ensure structures and procedures are in place to achieve the successful delivery of taxonomy engagements to time, to budget and to satisfactory quality standards.&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Full Job Description online at:&lt;br /&gt;
&lt;br /&gt;
http://careers.peopleclick.com/jobposts/Client40_DowJones/BU1/External/pck314-6888.htm&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Dave Clarke&amp;lt;br&amp;gt;&lt;br /&gt;
Global Taxonomy Director&amp;lt;br&amp;gt;&lt;br /&gt;
Dow Jones&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Start-up Seeking Software Engineer Computational Linguistics  ==&lt;br /&gt;
&lt;br /&gt;
We are a data driven start-up looking for a CTO to build a social media analytics application&lt;br /&gt;
using Natural Language Processing with experience in web-scraping and text mining.&lt;br /&gt;
Candidate will develop methodologies to find sentiment&lt;br /&gt;
(sentiment in context, polarity, intensity, relationships between issues and reasons for issues),&lt;br /&gt;
tuning for less formal content (tokenization, sentence segmentation, part-of-speech (POS) tagging, parsing, etc.)&lt;br /&gt;
and finding embedded meaning and intelligence in less grammatical text.&lt;br /&gt;
Experience in predictive modeling in data mining a plus.&lt;br /&gt;
&lt;br /&gt;
Qualifications:&lt;br /&gt;
* 2-4 years experience in a relevant field such as software engineering, machine learning programming, et al.&lt;br /&gt;
* The drive to work in a fast-paced, multi-disciplinary start-up environment.&lt;br /&gt;
* A strong interest in online social networks and network/user dynamics.&lt;br /&gt;
* PhD in a computational linguistics (or related field) required.&lt;br /&gt;
* Experience in database design and search algorithms.&lt;br /&gt;
* Experience with Python/MySQL/R and SaaS is preferred but not required.&lt;br /&gt;
* Excellent documentation &amp;amp; whitepaper authoring skills.&lt;br /&gt;
* Patent filing experience preferred.&lt;br /&gt;
&lt;br /&gt;
I am looking for an exceptional management team member who will provide design the product vision&lt;br /&gt;
and has extensive experience in delivering complex web-based applications of significant scale&lt;br /&gt;
versed in the social semantic web.&lt;br /&gt;
Feedback of the product framework has been extremely positive;&lt;br /&gt;
we are looking to build a working prototype of site.&lt;br /&gt;
There is no direct compensation for this role.&lt;br /&gt;
Equity stake in company for right candidate.&lt;br /&gt;
This role will become a full time position once funded.&lt;br /&gt;
&lt;br /&gt;
This position will require telecommuting as we are located in Newburyport, MA.&lt;br /&gt;
&lt;br /&gt;
Learn more: http://www.socialtality.com &lt;br /&gt;
&lt;br /&gt;
Contact me directly at wendytroupe [at] socialtality [dot] com&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Semantic_Web_Jobs&amp;diff=6652</id>
		<title>Semantic Web Jobs</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Semantic_Web_Jobs&amp;diff=6652"/>
		<updated>2026-06-20T07:40:52Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;If you&#039;d like to hire a Semantic Web expert or if you&#039;re looking for a Semantic Web position,&lt;br /&gt;
please send me a short note or an HTTP link.&lt;br /&gt;
If you&#039;re posting a position, please be sure to note whether or not the job is in the NYC area.&lt;br /&gt;
If you&#039;re looking for a position, please indicate whether you&#039;re willing to consider locations besides the NYC area.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Vice President- Ontologist JPMorgan Chase NYC or NJ===&lt;br /&gt;
&lt;br /&gt;
https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/job/210760263&lt;br /&gt;
&lt;br /&gt;
6/20/2026&lt;br /&gt;
&lt;br /&gt;
The Firmwide Chief Data Office is responsible for maximizing the value and impact of data globally, in a highly governed way. It consists of several teams focused on accelerating JPMorgan Chase’s data, analytics, and AI journey, including data strategy, data impact optimization, privacy, data governance, transformation, and talent. We are looking for a data executive to help shape the strategy for how we make data available to power everything from new product development to Artificial Intelligence models. This leader will join the team responsible for setting the firmwide data publishing strategy and driving the adoption of the strategy across the firm.&lt;br /&gt;
&lt;br /&gt;
As a Vice President-Ontologist within the JP Morgan Chase team, you will be instrumental in shaping our knowledge representation. Your role will involve utilizing ontologies and taxonomies to enhance data interoperability and management, preparing our data for AI applications. Your responsibilities will range from engaging with and educating domain experts, to assessing standard ontologies and developing our organization-wide ontology. Your work will traverse multiple domains, influencing departments like Data &amp;amp; Analytics, Product, and Tech.&lt;br /&gt;
&lt;br /&gt;
Job responsibilities:&lt;br /&gt;
&lt;br /&gt;
Development and adoption of ontologies to represent complex domains&lt;br /&gt;
&lt;br /&gt;
Evaluate industry standard ontologies for adoption across JPMC&lt;br /&gt;
&lt;br /&gt;
Work closely with stakeholders, subject matter experts, product owners, and engineers to understand their use cases, requirements, and dependencies, critically assessing proposed solutions&lt;br /&gt;
&lt;br /&gt;
Provide expert input into the JP Morgan Chase’s firmwide data strategy&lt;br /&gt;
&lt;br /&gt;
Communicate complex ideas effectively to collaborators using precise terminology and relatable examples, and ask clarifying questions to define core meanings.&lt;br /&gt;
&lt;br /&gt;
Mentor fellow ontologists to ensure alignment with accepted practices, standards, objectives, key results, and strategic initiatives.&lt;br /&gt;
&lt;br /&gt;
Keep abreast of emerging trends and advancements in ontology engineering, knowledge representation, and semantic technologies.&lt;br /&gt;
&lt;br /&gt;
Balance timeliness with quality under tight deadlines, managing multiple priorities and partners.&lt;br /&gt;
&lt;br /&gt;
Ensure end-to-end relevance to stakeholder needs, from gathering competency questions to achieving successful integrations.&lt;br /&gt;
&lt;br /&gt;
Required qualifications, capabilities, and skills:&lt;br /&gt;
&lt;br /&gt;
3+ years of experience developing and managing ontologies for real-world applications&lt;br /&gt;
&lt;br /&gt;
Expertise in Data and Financial service standards such as ISO 20022, OWL, RDF, SKOS, and SHACL.&lt;br /&gt;
&lt;br /&gt;
Experience with ontology and taxonomy development process and tools (e.g., Protégé, TopBraid Composer, PoolParty, etc.).&lt;br /&gt;
&lt;br /&gt;
Structured thinker and effective communicator with excellent written communication skills. Ability to crisply articulate complex technical concepts to senior audiences with poise and confidence.&lt;br /&gt;
&lt;br /&gt;
Preferred qualifications, capabilities, and skills:&lt;br /&gt;
&lt;br /&gt;
Master&#039;s or Ph.D. in a field focused on ontology engineering, knowledge representation, or semantic technologies, such as Information Science, Library Science, Philosophy, Linguistics, or Computer Science.&lt;br /&gt;
&lt;br /&gt;
Experience with Financial sector data standards and ontologies&lt;br /&gt;
&lt;br /&gt;
Experience with program management and collaborative development best practices&lt;br /&gt;
&lt;br /&gt;
Understanding of large-scale, distributed, end-to-end systems.&lt;br /&gt;
&lt;br /&gt;
Knowledge of Data Governance and Data Management&lt;br /&gt;
&lt;br /&gt;
Contributions to the Ontology community, such as papers, conference presentations, industry standard contributions, or open source contributions (e.g., Github repos)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Senior Lead Software Engineer-Ontology and RDF - JPMorgan Chase===&lt;br /&gt;
&lt;br /&gt;
Location:  GLASGOW, LANARKSHIRE, United Kingdom &lt;br /&gt;
&lt;br /&gt;
As a Lead Software Engineer at JPMorgan Chase within the Identity and Access Management, Corporate Sector, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.&lt;br /&gt;
&lt;br /&gt;
https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/requisitions/preview/210472446/?keyword=Ontology&lt;br /&gt;
&lt;br /&gt;
===Lead Software Engineer- Ontology and RDF - JPMorgan Chase===&lt;br /&gt;
&lt;br /&gt;
Location:  GLASGOW, LANARKSHIRE, United Kingdom &lt;br /&gt;
&lt;br /&gt;
As a Lead Software Engineer at JPMorgan Chase within the Identity and Access Management , Corporate Sector, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.&lt;br /&gt;
&lt;br /&gt;
https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/requisitions/preview/210472450/?keyword=Ontology&lt;br /&gt;
&lt;br /&gt;
===Ontologist IMDb Bristol===&lt;br /&gt;
&lt;br /&gt;
URL: Allocated &amp;lt;!-- https://www.amazon.jobs/en-gb/jobs/2026698/ontologist-imdb-content --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Location: Bristol, UK&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
DESCRIPTION&lt;br /&gt;
Job summary&lt;br /&gt;
With more than 400 million searchable data items — including 10 million movie, TV and entertainment titles, 11 million cast and crew members and 11 million images — IMDb is the world’s most popular and authoritative source for information on movies, TV shows and celebrities, and has a combined web and mobile audience of more than 200 million monthly visitors. The IMDb database is continually growing, thanks to a vast contributor community of entertainment professionals and companies, IMDb staff, individual contributors and other trusted sources. IMDb content is integrated into strategically important parts of Amazon and AWS businesses, including Amazon Fire TV, Alexa, and X-Ray on Prime Video. IMDb licenses information from its vast and authoritative database to third-party businesses, including film studios, television networks, streaming services and cable companies, as well as airlines, electronics manufacturers, non-profit organizations and software developers. Learn more at developer.imdb.com. Other IMDb products and services include: the IMDb website for desktop and mobile devices; apps for iOS and Android; a free streaming channel, IMDb TV; and IMDb original video series and podcasts. For entertainment industry professionals, IMDb provides IMDbPro and Box Office Mojo. IMDb is an Amazon company. For more information, visit imdb.com/press and follow @IMDb.&lt;br /&gt;
&lt;br /&gt;
IMDb is a group of entertainment enthusiasts – and we are passionate about ensuring all our customers around the globe have access to all the content they need, when they need it. IMDb sits at the intersection of the entertainment, media, and technology markets inside the world’s most innovative and consumer-centric company – Amazon.com. IMDb employees enjoy the benefits of working for Amazon with the autonomy of working on a smaller, nimble team.&lt;br /&gt;
&lt;br /&gt;
As an Ontologist, you work as part of a global team to deliver world-class, intuitive, and comprehensive taxonomy and ontology models to optimize product delivery for IMDb web and mobile experiences. You collaborate with business partners and engineering teams to deliver knowledge-based solutions to enable discovery and engagement with IMDb content. In this role you will directly impact the customer experience as well as the company&#039;s product knowledge foundation across all customer cohorts.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Specific responsibilities include the following:&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Develop logical, semantically rich, and extensible data models for IMDb&#039;s expansive entertainment catalog.&lt;br /&gt;
&lt;br /&gt;
Ensure our ontologies provide comprehensive sub-domain coverage that are available for machine ingestion and inference&lt;br /&gt;
&lt;br /&gt;
Research worldwide understanding of entertainment content to develop scalable data models that solve customer problems and enhance entity discovery&lt;br /&gt;
&lt;br /&gt;
Contribute to the development of new tools, features and processes for the Ontology team&lt;br /&gt;
&lt;br /&gt;
Support the content expansion team focused on driving the overall IMDb Content product and business strategy and execution.&lt;br /&gt;
&lt;br /&gt;
BASIC QUALIFICATIONS&lt;br /&gt;
*Experience working in ontology and/or taxonomy roles&lt;br /&gt;
*Proven skills in data retrieval and data research techniques&lt;br /&gt;
*Ability to quickly understand complex processes and communicate them in simple language&lt;br /&gt;
*Ability to communicate knowledge-based requirements and needs to engineering and retail teams&lt;br /&gt;
*Familiarity with Semantic Web technologies (RDF/s, OWL), query languages (SPARQL) and validation/reasoning standards (SHACL, SPIN)&lt;br /&gt;
*Detail-oriented problem-solving, ability to work in fast-changing environment and manage ambiguity&lt;br /&gt;
*Proven track record of strong communication and interpersonal skills&lt;br /&gt;
*Proficient English language skills&lt;br /&gt;
&lt;br /&gt;
PREFERRED QUALIFICATIONS&lt;br /&gt;
Master’s degree in Library Science, Information Systems, Linguistics or other relevant fields&lt;br /&gt;
&lt;br /&gt;
Experience building ontologies in the entertainment and semantic search spaces&lt;br /&gt;
&lt;br /&gt;
Experience working with schema-level constructs (e.g. higher-order classes, punning, property inheritance)&lt;br /&gt;
&lt;br /&gt;
Proficiency in SQL, SPARQL&lt;br /&gt;
&lt;br /&gt;
Familiarity with software engineering life cycle&lt;br /&gt;
&lt;br /&gt;
Familiarity with ontology manipulation programming libraries&lt;br /&gt;
&lt;br /&gt;
Exposure to data science and/or machine learning, including graph embeddings&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Amazon is an equal opportunities employer. We believe passionately that employing a diverse workforce is central to our success. We make recruiting decisions based on your experience and skills. We value your passion to discover, invent, simplify and build. Protecting your privacy and the security of your data is a longstanding top priority for Amazon. Please consult our Privacy Notice (https://www.amazon.jobs/en/privacy_page) to know more about how we collect, use and transfer the personal data of our candidates.&lt;br /&gt;
&lt;br /&gt;
==KNOWLEDGE GRAPH AND SEMANTICS SME==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;About the job&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Company: Solutions Driven &lt;br /&gt;
&lt;br /&gt;
Location: United States &lt;br /&gt;
&lt;br /&gt;
Type: Remote &lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;About the team:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Our client&#039;s Data Team helps to transform every aspect of their business. We are highly skilled at formulating data strategy, defining business and technology initiatives across the data management lifecycle, and aligning multi-year strategic roadmaps with client’s business goals. As digital technologies advance and regulations tighten, today’s consumers – and, therefore, today’s businesses – are becoming more aware of the importance of good quality data. We work to establish holistic ways to effectively manage data through the modern data supply chain and facilitate consumption through analytics, modelling, AI, machine learning, dashboarding, and reporting.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;What You’ll Get to Do:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The practice leader will involve shaping and executing the hiring strategy, developing the best of in-house talent, leading talent, and collaboration with the others in the data &amp;amp; analytics practice to bring cutting edge solutions to the market.&lt;br /&gt;
&lt;br /&gt;
Graph and Semantic Engineering are emerging enterprise data technologies that leverage ontologies and related advanced analytics – including artificial intelligence – to build and represent knowledge in expressive and novel ways&lt;br /&gt;
&lt;br /&gt;
Primary responsibility in this role is to help clients succeed by understanding, exploiting, designing, and implementing solutions involving graph and semantic technologies –  Solutions Driven United States Remote including advanced data management and analytics, and other innovative new ways&lt;br /&gt;
&lt;br /&gt;
In the fast-changing landscape of potential partners in graph technology, an ongoing curiosity and developing connections with key partners, will enable valuable support for clients&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;What You’ll Bring with You:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Graph and Semantic Data Consultants will require multi-disciplinary capabilities required to explain, design and create powerful knowledge graph models – a good balance of business and technical understanding&lt;br /&gt;
&lt;br /&gt;
Sound understanding of Knowledge Representation and Semantic Technologies (OWL, RDF, SWRL, SPARQL, JSON-LD) including semantic modelling and data integration, data unification, knowledge graph design, ontology and taxonomy&lt;br /&gt;
&lt;br /&gt;
Awareness of the breadth of use cases for these technologies, prudent user interface design which exploit the power of graph and mitigates the challenges of visualizing graph data&lt;br /&gt;
&lt;br /&gt;
Experience with Graph/Triple Stores and related technologies, e.g. PoolParty, Stardog, TigerGraph, Ontotext GraphDB, Grakn, MarkLogic, Metaphactory, Neo4j&lt;br /&gt;
&lt;br /&gt;
Understanding of data/information architecture and modeling languages such as UML, ER, IDEF, Data Flow Diagram&lt;br /&gt;
&lt;br /&gt;
Excellent Communication skills – be an effective, passionate, trusted advocate and communicator for Graph and Semantic Technologies&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;Other desired skills:&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Good awareness of financial services domain is a great advantage&lt;br /&gt;
* Ability to utilize NLP, machine learning and semantic text mining to translate data into machine understandable representations to support advanced analytics, classification, inference processing and knowledge extraction&lt;br /&gt;
* Familiarity with mapping techniques, e.g. R2RML, for transforming relational/tabular datasets into triples&lt;br /&gt;
* Agile software development frameworks, e.g. Scrum, Kanban, Extreme Programming&lt;br /&gt;
* Understanding of data management and governance toolsets and methodologies&lt;br /&gt;
* Experience in applying semantic technologies in an industry setting, or practical experience as part of an academic, Library Science or Information/Knowledge &lt;br /&gt;
* Management career &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;Why Us?&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We are the largest Financial Services focused consultancy in the world, serving everyone from global banks to emerging FinTechs, from strategy through digital transformation, design, business consulting, data and analytics, cyber, cloud, technology architecture, and engineering. We are young and growing firm. We maintain an entrepreneurial spirit and growth mindset, and have minimal bureaucracy. We have no internal silos that get in the way of your career opportunities or ability to focus on our clients and make a difference to the business. We offer the opportunity for everyone to learn rapidly, take on tough challenges, and get promoted quickly. We take pride in our creative, collaborative, diverse, and inclusive culture, where everyone can Be Yourself at Work. We offer highly competitive benefits, including medical, dental and vision insurance, a 401(k) plan, tuition reimbursement, and a work culture focused on innovation and creation of lasting value for our clients and employees.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;b&amp;gt;Ready to take the Next Step&amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If this sounds like you, we would love to hear from you. This is an opportunity to make a difference and contribute to a highly successful company with a significant growth trajectory.&lt;br /&gt;
&lt;br /&gt;
This position may be performed remotely anywhere within the United States except the State of Colorado.&lt;br /&gt;
&lt;br /&gt;
Apply here: https://www.linkedin.com/jobs/view/2943587552/&lt;br /&gt;
&lt;br /&gt;
==Semantics Specialist==&lt;br /&gt;
&lt;br /&gt;
Company: [A]&lt;br /&gt;
&lt;br /&gt;
Location: 100% Remote from North America (&#039;&#039;&#039;excluded:&#039;&#039;&#039; CA, MA or NY)&lt;br /&gt;
&lt;br /&gt;
[A] is looking for a Semantics Specialist to help make the world smarter with intelligent content as a part of the [A] team. [A] seeks motivated, creative, innovative candidates eager to start a client-facing role with a mix of taxonomy, thesaurus, and ontology development experience with an interest in learning or growing their experience in content development and content modeling. &lt;br /&gt;
&lt;br /&gt;
The ideal candidate will have experience and knowledge working on the following content principles:&lt;br /&gt;
&lt;br /&gt;
Content Semantics (controlled vocabularies, ontology, taxonomy, metadata, schemas etc.). Content Structure (content modeling, markup, schema, DITA, XML) This role is NOT focused on Content &amp;quot;Interchange&amp;quot; (metadata, XSLT, XML Processors and Parsers), or Content Operations (CMS tools and platforms management, Content acquisition tools, systems, processes, and patterns). At [A], these practices have their own speciality focus. Having familiarity with these related practices is very helpful. &lt;br /&gt;
&lt;br /&gt;
The responsibilities for this role deeply involve the principles of content semantics and structure. And this role requires one to build an understanding of the application of semantics principals at abstract, architecture, and at very tactical applied levels within enterprise knowledge and customer experience delivery environments.&lt;br /&gt;
&lt;br /&gt;
This is a remote job that can be done mostly from your home office anywhere in North America (we are unable to consider candidates in the states of CA, MA or NY at this time), but may involve occasional travel to a client location (once safe to do so). &lt;br /&gt;
&lt;br /&gt;
The Semantics Specialist will work within the larger Content Intelligence practice to build next-generation content systems for clients, helping them craft intelligent content for multi-channel marketing and increase the value of semantically-rich content. You will participate in, or lead, client projects that include development of semantic models, taxonomies, thesauri, controlled vocabularies, and formal ontologies in conjunction with content structures and content tools experts.&lt;br /&gt;
&lt;br /&gt;
You’ll define content relationships by using taxonomies, ontologies, and metadata. This includes implementation of semantic models within semantic software tools such as those from Synaptica, the Semantic Web Company, and other semantic provider solutions that are fitting into new Content Intelligence architectures.&lt;br /&gt;
&lt;br /&gt;
Link: http://jobs.simplea.com/apply/cTn4Ye5FlK/Semantics-Specialist&lt;br /&gt;
&lt;br /&gt;
==Head of Ontologies and Business Domain Data Modeling - NJ USA==&lt;br /&gt;
&lt;br /&gt;
Date: 09/2020&lt;br /&gt;
&lt;br /&gt;
Location: East Hanover, New Jersey, USA&lt;br /&gt;
&lt;br /&gt;
Strategic purpose of this role is to organize Novartis data, make it easily accessible, and useful to authorized roles. This role will help drive the development and execution of Novatis’ ambition to turn data into a real strategic asset across the organization. This ambition is one of key pillars in the broader digital transformation happening at Novartis to be a ‘medicines and data science company.’ This data centric role will facilitate using data to digitize the biopharma value chain in finding right target, right tissue, right safety, right patients, right commercial potential and to optimize omni channel stakeholder experience and engagement. More specifically, the purpose of this role is to engineer, in partnership with over 10+ business units and Technology partners (internal and external), global adoption of consistent Data Life Cycle (DLC) management framework, processes, architecture and supporting tools. Key areas in the Data Life Cycle management includes Master Data Management, Data Governance, Information Modeling, and Data Science Enablement. This role will specifically focus on improving maturity of Data Ownership and Capability Building in all bio pharma data domains across the value chain; omics, compounds, diseases, patients, payers, providers, sites, trials, KOLs, HCPs, EMR, EHR, patient journeys, employees, contracts, vendors, products, etc&lt;br /&gt;
&lt;br /&gt;
This role will directly report and work with the Head of Data Strategy in the Group Digital Office.&lt;br /&gt;
&lt;br /&gt;
URL: https://careers.iscb.org/jobs/view/7251&lt;br /&gt;
&lt;br /&gt;
==Head of Knowledge Graph &amp;amp; Semantics - NJ USA==&lt;br /&gt;
&lt;br /&gt;
Date: 09/2020&lt;br /&gt;
&lt;br /&gt;
Location: East Hanover, New Jersey, USA&lt;br /&gt;
&lt;br /&gt;
Strategic purpose of this role is to organize Novartis data, make it easily accessible, and useful to authorized roles. This role will help drive the development and execution of Novartis’ ambition to turn data into a real strategic asset across the organization. This ambition is one of key pillars in the broader digital transformation happening at Novartis to be a ‘medicines and data science company.’ This data centric role will facilitate using data to digitize the biopharma value chain in finding right target, right tissue, right safety, right patients, and right commercial potential and to optimize Omni channel stakeholder experience and engagement.&lt;br /&gt;
&lt;br /&gt;
More specifically, the purpose of this role is to drive, in partnership Digital Data Science and AI teams, over 10+ business units, Technology partners (internal and external) adoption of consistent framework, processes, architecture and supporting tools for analytics needs. This role will help create high quality data assets to enable analytics using data across all bio pharma data domains and value chain; omics, compounds, diseases, patients, payers, providers, sites, trials, KOLs, HCPs, EMR, EHR, patient journeys, employees, contracts, vendors, products, etc.&lt;br /&gt;
&lt;br /&gt;
This role will directly report and work with the Head of Data Strategy in the Group Digital Office.&lt;br /&gt;
&lt;br /&gt;
URL: https://careers.iscb.org/jobs/view/7250&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Software Engineer to work on the Data Team at Financial Services Firm NYC/NJ==&lt;br /&gt;
Date: 08/2020&lt;br /&gt;
&lt;br /&gt;
Location: NYC /NJ&lt;br /&gt;
&lt;br /&gt;
Software Engineer to work on the data team for a premier financial services firm in NYC/NJ on Enterprise Data/Historical Data/Metadata / Semantic Wikis/SPARQL/OWL – Ontology Semantic Web Language/RDFS.&lt;br /&gt;
&lt;br /&gt;
status: filled&lt;br /&gt;
&lt;br /&gt;
==Analytical Linguist, Knowledge Engine Google - San Francisco, CA, USA ==&lt;br /&gt;
Behind Google&#039;s Knowledge Engine (KE) team is the world&#039;s largest and most comprehensive semantic graph. We power features across Google, including Ads, YouTube, Google Play, Geo, and many others. Growing Knowledge Graph requires precise models describing how entities fit together. As an Analytical Linguist on the Knowledge Engine team, you will analyze graph structures and content, develop new semantic representations, and work with providers and consumers of data to guide the development and usage of knowledge structures. Using a variety of semantic modeling techniques, you’ll work to improve entity accessibility. You will be responsible for judging tradeoffs between formality and usability and deciding when to borrow from an existing ontology and when to develop new structures.&lt;br /&gt;
&lt;br /&gt;
Responsibilities&lt;br /&gt;
Analyze graph structures and content.&lt;br /&gt;
Develop new semantic representations.&lt;br /&gt;
Work with providers and consumers of data to guide the development and usage of knowledge structures to improve entity accessibility.&lt;br /&gt;
Use a variety of sematic modeling techniques, judging tradeoffs between formality and usability and deciding when to borrow from an existing ontology and when to develop new structures.&lt;br /&gt;
&lt;br /&gt;
Minimum qualifications&lt;br /&gt;
MA/MS in linguistics, computer science, library/information science, philosophy, or a related field, or equivalent work experience.&lt;br /&gt;
Coding experience in Python, C/C++, Java, or Go.&lt;br /&gt;
Preferred qualifications&lt;br /&gt;
PhD in relevant field, or equivalent work experience.&lt;br /&gt;
Experience with ontology development (RDF/OWL, SPARQL, the Semantic Web or Frame based KR systems).&lt;br /&gt;
Area&lt;br /&gt;
There is always more information out there, and the Knowledge team has a never-ending quest to find it and make it accessible. We&#039;re constantly refining our signature search engine to provide better results, and developing offerings like Google Instant, Google Voice Search and Google Image Search to make it faster and more engaging. We&#039;re providing users around the world with great search results every day, but at Google, great just isn&#039;t good enough. We&#039;re just getting started.&lt;br /&gt;
&lt;br /&gt;
Team or role:Program Management&lt;br /&gt;
Job type:Full-time&lt;br /&gt;
Last updated: Dec 31, 2014&lt;br /&gt;
Job location(s):San Francisco, CA, USA&lt;br /&gt;
&lt;br /&gt;
https://www.google.com/about/careers/search#!t=jo&amp;amp;jid=86625001&amp;amp;&lt;br /&gt;
&lt;br /&gt;
==Wiss. Mitarbeiterin / Wiss. Mitarbeiter befristet auf 6 Monate (E 13 TV-L FU)==&lt;br /&gt;
&lt;br /&gt;
Im Rahmen der Juniorprofessur für Web Science/Human-Centered Computing (Prof. Dr. Claudia Müller-Birn) an der FU Berlin wurde im letzten Jahr ein Team aufgebaut, das schwerpunktmäßig in den Bereichen der Web-basierten Wissensgenerierung und des Digitalen Lernens tätig ist. Die Gruppe arbeitet mit Universitäten und Forschungseinrichtungen sowie non-profit Organisationen national und international zusammen.&lt;br /&gt;
&lt;br /&gt;
Weitere Informationen unter: http://www.mi.fu-berlin.de/inf/groups/hcc/index.html&lt;br /&gt;
&lt;br /&gt;
Das Aufgabengebiet umfasst die Mitarbeit in dem Forschungsprojekt Neonion – Collaborative Annotation in Humanities and Arts –, innerhalb dessen eine semantische Annotationssoftware in einem nutzerzentrierten Designprozesses weiterentwickelt wird. In einem Pilotprojekt mit einem Projektpartner aus dem Bereich der Geisteswissenschaften werden neuartige Interaktionskonzepte zur semantischen Annotation von Dokumenten erforscht. Eine aktive Beteiligung an laufenden Forschungsanträgen und Publikationsaktivitäten wird erwartet. Eine Möglichkeit zur Verlängerung dieser zeitlich befristeten Tätigkeit ist je nach Verfügbarkeit einer Folgefinanzierung gegeben.&lt;br /&gt;
&lt;br /&gt;
Einstellungsvoraussetzung ist ein abgeschlossenes Hochschulstudium in Informatik oder Wirtschaftsinformatik (Diplom oder Master). Alternativ ist ein Abschluss in einem verwandten Bereich mit entsprechender Erfahrung in der Informatik möglich.&lt;br /&gt;
&lt;br /&gt;
Folgende Kenntnisse sind für die Besetzung der Stelle von Vorteil und wünschenswert:&lt;br /&gt;
• Kenntnisse in den Bereichen der Semantischen Technologien und Web Technologien&lt;br /&gt;
• Kenntnisse im Bereich der nutzerzentrierten Softwareentwicklung und Usability&lt;br /&gt;
• Erfahrung im wissenschaftlichen Arbeiten und insbesondere in qualitativen Forschungsmethoden&lt;br /&gt;
• Kommunikationsfähigkeit und eigenständiges Arbeiten&lt;br /&gt;
• Erfahrungen im Umgang mit Confluence and Jira&lt;br /&gt;
• Bereitschaft zur interdisziplinären Zusammenarbeit in einem internationalen Konsortium&lt;br /&gt;
• Sehr gute Englische Kenntnisse in Wort und Schrift&lt;br /&gt;
&lt;br /&gt;
Bitte senden Sie Ihre Bewerbung per E-Mail bis zum 30.06.2014 als eine einzige PDF-Datei mit einem ausführlichen Lebenslauf (inkl. Kenntnisse und praktischen Erfahrungen), Kopien Ihrer akademischen Abschlüsse, Referenzschreiben (falls vorhanden) und ein Motivationsschreiben an Prof. Dr. Cl. Müller-Birn (clmb@inf.fu-berlin.de).&lt;br /&gt;
&lt;br /&gt;
Für weitere Informationen zu dieser Ausschreibung wenden Sie sich bitte ebenfalls per E-Mail an Prof. Dr. Cl. Müller-Birn (clmb@inf.fu-berlin.de).&lt;br /&gt;
&lt;br /&gt;
==Wissenschaftliche/r Mitarbeiter/in (Data Scientist) am Museum für Naturkunde Berlin (MfN)==&lt;br /&gt;
&lt;br /&gt;
http://www.naturkundemuseum-berlin.de/fileadmin/startseite/institution/stellenausschreibungen/17_2014.pdf&lt;br /&gt;
&lt;br /&gt;
Arbeitszeit: 100 v.H. d. regelm. wöchentlichen Arbeitszeit&lt;br /&gt;
&lt;br /&gt;
Befristung: Drittmittelfinanzierung, zum nächstmöglichen Zeitpunkt befristet bis zum 31. Mai 2015 &lt;br /&gt;
&lt;br /&gt;
Entgeltgruppe: E13 TV-L Berlin &lt;br /&gt;
&lt;br /&gt;
Aufgabengebiet: Die erfolgreiche Bewerberin / der erfolgreiche Bewerber soll das Museum für Naturkunde in zwei Projekten des Forschungsbereichs „Digitale Welt und Informationsmanagement“ unterstützen: Das EU-geförderte Projekt EuropeanaCreative (www.pro.europeana.eu/web/europeana-creative) befasst sich mit den Möglichkeiten der Nachnutzung digitaler Medien über die Plattform Europeana (www.europeana.eu). Das DFG-geförderte Projekt „German Federation for the Curation of Biological Data” (www.gfbio.org) arbeitet am Aufbau einer grundlegenden Forschungsinfrastruktur, die den Austausch von biologischen und umweltbezogenen Forschungsdaten ermöglichen bzw. vereinfachen soll. Nach der derzeitigen 18-monatigen Pilotphase sind zwei je 36 Monate laufende Anschlussprojekte geplant. &lt;br /&gt;
&lt;br /&gt;
Die Stelle eignet sich als Startpunkt für den Aufbau einer Forschergruppe am MfN mit Fokus auf semantischer Biodiversitätsinformatik. Das Ziel des Forschungsbereichs „Digitale Welt und Informationsmanagement“ ist die Implementierung einer zukunftsfähigen Architektur für naturkundliche Daten und Datenrepositorien, welche neue Möglichkeiten zur Beantwortung von Forschungsfragen sowohl im Bereich „Data Science“ als auch „Public Engagement with Science“ eröffnet.&lt;br /&gt;
&lt;br /&gt;
Der Aufgabenbereich der erfolgreichen Bewerberin / des erfolgreichen Bewerbers umfasst:&lt;br /&gt;
* Methoden und Prozesse von Linked Open Data, semantischem Web und Wissensmanagement in Hinblick auf ihre Nutzbarkeit für naturkundliche * Daten zu evaluieren und exemplarisch zu testen bzw. umzusetzen&lt;br /&gt;
* Grundlagen für eine effektivere Zusammenarbeit von Forschern verschiedener Fachrichtungen, z.B. Taxonomie, Umweltwissenschaften und  Naturschutz, mittels semantischen Wissensmanagements zu erarbeiten&lt;br /&gt;
* Möglichkeiten von „Data-driven-science“ und „Open Science“ in Pilotprojekten zu demonstrieren&lt;br /&gt;
* konzeptionelle und strategische Weiterentwicklung von GFBio in Zusammenarbeit mit dem GFBio Team am MfN voranzutreiben&lt;br /&gt;
* in Kooperation mit GFBio Projektpartnern am Aufbau des GFBio Terminologieservers mitzuwirken und am MfN vorhandene Terminologien bzw. Ontologien für die Verwendung in GFBio zu erschließen&lt;br /&gt;
* Erschließung, Analyse und Impact naturkundlicher Medien von Europeana Creative mit Mitteln des Semantic Webs zu verbessern&lt;br /&gt;
* Unterstützung der Vermittlung der Projektergebnisse nach innen und außen durch Kommunikation, Präsentationen, Berichte und Publikationen&lt;br /&gt;
&lt;br /&gt;
Anforderungen:&lt;br /&gt;
* abgeschlossenes wissenschaftliches Hochschulstudium der Informatik, Bioinformatik oder verwandter Fächer&lt;br /&gt;
* Nachweis aktiver wissenschaftlicher Publikationstätigkeit&lt;br /&gt;
* praktische Erfahrungen im Bereich Linked Open Data, Semantic Web oder Ontologieentwicklung&lt;br /&gt;
* sehr gute Kenntnisse der deutschen und englischen Sprache in Wort und Schrift&lt;br /&gt;
* ausgezeichnete Kommunikations- und Teamfähigkeit, insbesondere in interdisziplinären Teams&lt;br /&gt;
* selbstständiges und strukturiertes Arbeiten&lt;br /&gt;
* Freude an der Erprobung innovativer Vermittlungswege für wissenschaftliche bzw. naturkundliche Fragestellungen&lt;br /&gt;
* überdurchschnittliche Einsatzbereitschaft und Zuverlässigkeit&lt;br /&gt;
&lt;br /&gt;
Von Vorteil sind:&lt;br /&gt;
* Promotion&lt;br /&gt;
*Erfahrung im Informations- und Wissensmanagement an wissenschaftlichen Einrichtungen&lt;br /&gt;
*Erfahrung in der Anleitung von Teams&lt;br /&gt;
*Erfahrungen mit Semantic Media Wiki oder Wikidata&lt;br /&gt;
*Kenntnisse der Aufgaben und Arbeitsweisen naturkundlicher Forschungsmuseen&lt;br /&gt;
&lt;br /&gt;
==Data Analytics Technical Architect - Chase Auto Finance - Jersey City, NJ==&lt;br /&gt;
Chase is a leader in the financial services industry, providing banking, mortgages, credit cards, loans, payment processing and investment services to 50 million customers - 1 out of every 6 Americans. As a division of JPMorgan Chase &amp;amp; Co. (NYSE:JPM), we:&lt;br /&gt;
Serve 21 million households with consumer banking relationships&lt;br /&gt;
Lent $17 billion to small businesses in 2011&lt;br /&gt;
Are one of the nation&#039;s largest credit card issuers, with more than 64 million credit cards in circulation&lt;br /&gt;
Service 8 million mortgage and home equity loans&lt;br /&gt;
While we operate across a broad range of businesses, our mission at Chase is quite simple: to be the industry leader in customer service. Our employees put the firm&#039;s resources to work every day for our customers.&lt;br /&gt;
 &lt;br /&gt;
Chase offers a dynamic environment and the training and support to meet your full potential. Our company is widely recognized as a great place to work, to grow and to invest for the future. Join our team.&lt;br /&gt;
The Chase Auto Finance (CAF) Data Management Program is a multi-year technology initiative to support CAF and enterprise data management processes addressing operational, financial, regulatory and analytical reporting requirements.  The program includes the long term development of a strategic platform for CAF with data sourcing, data enrichment, analytical, monitoring and reporting capabilities.  This platform will also integrate with cross-LOB efforts to create a single consolidated customer view.&lt;br /&gt;
 &lt;br /&gt;
The CAF Data Management team is looking for a highly motivated individual that will have a strong foundational knowledge and experience with distributed systems and computing systems with hands-on engineering skills. Experience with a range of big data architectures and broad understanding and experience of real-time analytics, NoSQL data stores, big data analytics products. In addition, experience with data modeling and data management, analytical tools, languages, or libraries is required.&lt;br /&gt;
The role is heavily collaborative, working closely with several internal business units as well as a firm-wide organization driving standards and best practices and the candidate will have several years experience working in client-focused roles.&lt;br /&gt;
 &lt;br /&gt;
Key Areas of Responsibility:&lt;br /&gt;
This individual will be responsible for guiding the full lifecycle of a Big Data Analytics solution, including requirements analysis, platform selection, technical architecture design, application design and development, testing, and deployment.  We are looking for candidates with a broad set of technology skills to be able to design and build robust solutions for big data problems and learn quickly as the platform grows.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Hands-on experience with the Hadoop stack (e.g. MapReduce, Sqoop, Pig, Hive, Hbase, Flume)&lt;br /&gt;
Hands-on experience with related/complementary open source software platforms and languages (e.g. Java, Linux, Apache, Perl/Python/PHP, Chef)&lt;br /&gt;
Hands-on experience with ETL (Extract-Transform-Load) tools (e.g Informatica,  Talend, Pentaho)&lt;br /&gt;
Hands-on experience with analytical tools, languages, or libraries (e.g. SAS, SPSS, R, Mahout)&lt;br /&gt;
Hands-on experience with &amp;quot;productionalizing&amp;quot; Hadoop applications (e.g. administration, configuration management, monitoring, debugging, and performance tuning)&lt;br /&gt;
Previous experience with high-scale or distributed RDBMS (Teradata, Netezza, Greenplum, Aster Data, Vertica)&lt;br /&gt;
Knowledge of NoSQL platforms (e.g. key-value stores, graph databases, RDF triple stores)&lt;br /&gt;
Experience with at least 4 of the following activities in the context of high-scale or distributed systems:&amp;lt;br&amp;gt;&lt;br /&gt;
1.    Implementation of ETL applications&amp;lt;br&amp;gt;&lt;br /&gt;
2.    Implementation of reporting applications&amp;lt;br&amp;gt;&lt;br /&gt;
3.    Application/implementation of custom analytics&amp;lt;br&amp;gt;&lt;br /&gt;
5.    Data migration from existing data stores&amp;lt;br&amp;gt;&lt;br /&gt;
6.    Infrastructure and storage design&amp;lt;br&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
Bachelors Degree&lt;br /&gt;
Minimum 6 years systems development and implementation experience&lt;br /&gt;
Minimum 5+ years of hands on Java experience building scalable solutions&lt;br /&gt;
Minimum 2 years of hands-on experience with programming on high-scale or distributed systems (such as Hadoop)&lt;br /&gt;
Minimum 3+ years experience with full application development lifecycle (analysis, architecture, design, development, testing and deployment).&lt;br /&gt;
Minimum 3+ years experience working with multiple RDBMS and associated applications&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Preferred Skills&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Knowledge of the Auto Loan Industry&lt;br /&gt;
Experience in designing and implementation of Enterprise Information Integration architectures&lt;br /&gt;
An understanding of Business Intelligence concepts including Dashboards, and the data requirements of senior executives&lt;br /&gt;
Knowledge of Teradata Aster&lt;br /&gt;
Knowledge of Cognos&lt;br /&gt;
&lt;br /&gt;
JPMorgan Chase is an Equal Opportunity and Affirmative Action Employer, M/F/D/V&lt;br /&gt;
&lt;br /&gt;
https://jpmchase.taleo.net/careersection/2/jobdetail.ftl?lang=en&amp;amp;job=1239446&amp;amp;src=JB-13027&lt;br /&gt;
&lt;br /&gt;
==Senior Taxonomist in Financial Services New York, New York==&lt;br /&gt;
&lt;br /&gt;
Role Description:&lt;br /&gt;
&lt;br /&gt;
This role requires someone with a strong background in system and business analysis.  Key requirement of this position is strong communication skills both written and verbal, as the successful candidate will be expected to deal directly with business stakeholders to elicit requirements and clearly present these to the technical team for development.  Demonstration of requirements gathering, modelling and documentation would be required. &lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Responsibilities Include:&lt;br /&gt;
&lt;br /&gt;
- Developing, mapping, evaluating, and/or maintaining taxonomies&lt;br /&gt;
&lt;br /&gt;
- Conducting content audits and content analysis&lt;br /&gt;
&lt;br /&gt;
- Developing metadata schemas, tagging, conducting metadata audits&lt;br /&gt;
&lt;br /&gt;
- Conducting user research, usability testing Interviewing stakeholders, subject matter experts, and end users&lt;br /&gt;
&lt;br /&gt;
- Providing training on taxonomy maintenance, tagging, or tool use&lt;br /&gt;
&lt;br /&gt;
- Facilitating working sessions&lt;br /&gt;
&lt;br /&gt;
- Support taxonomy implementation&lt;br /&gt;
&lt;br /&gt;
- Support integration of content-, document-, taxonomy-management tools, search engines etc.&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Required Skills and Qualifications&lt;br /&gt;
&lt;br /&gt;
- Experience in working in document management space&lt;br /&gt;
&lt;br /&gt;
- Experience with--or knowledge of--taxonomy, metadata, controlled vocabularies, and classification&lt;br /&gt;
&lt;br /&gt;
- Excellent attention to detail&lt;br /&gt;
&lt;br /&gt;
-Very strong communication, both written and verbal.&lt;br /&gt;
&lt;br /&gt;
-Ability to lead/direct small focused groups of individuals.&lt;br /&gt;
&lt;br /&gt;
-Experience in dealing with key business stakeholders in a professional manner including client engagement and building relationships with business and partnering technology teams.&lt;br /&gt;
&lt;br /&gt;
contact: info@kona.llc&lt;br /&gt;
&lt;br /&gt;
==Software Engineer In The Semantic Algorithms And News Delivery==&lt;br /&gt;
Location:  Roseland, NJ&lt;br /&gt;
&lt;br /&gt;
Date: 29 September 2012&lt;br /&gt;
&lt;br /&gt;
Acquire Media analyzes and distributes the news. We have an opening for a software engineer to help our semantic algorithms and news delivery teams.&lt;br /&gt;
&lt;br /&gt;
Here&#039;s the link: (on Craigslist) http://newjersey.craigslist.org/sof/3409136469.html&lt;br /&gt;
&lt;br /&gt;
== Senior Java Developer With Extensive Semantic Web Experience - Nature Publishing Group ==&lt;br /&gt;
&lt;br /&gt;
Location: London, UK&amp;lt;br&amp;gt;&lt;br /&gt;
Date: 27 November 2012&lt;br /&gt;
&lt;br /&gt;
Nature Publishing Group - the prestigious international scientific publishing company - seeks an enthusiastic Software Developer to work&lt;br /&gt;
in our London office as an integral part of the team developing our growing portfolio of online products.&lt;br /&gt;
&lt;br /&gt;
You will enhance and expand products across http://www.nature.com, ensuring that your work is well designed, maintainable and tested. You will fit this description: http://bit.ly/hacker-def and enjoy having conversations about new frameworks, new technologies and the joys of distributed version control systems.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The successful candidate must:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
have great technical ability in the field of web development write high quality code delight in having an intimate understanding of the internal workings of a system be comfortable talking about design, code and the requirements that shape it enjoy the intellectual challenge of creatively overcoming or circumventing limitations have a great desire to learn, share knowledge and crave constructive criticism&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;You will also be able to demonstrate the required skills:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
extensive experience of Java 6, dependency injections frameworks and relational databases experience in web servers (Jetty, Tomcat), XML, web services, JUnit, Mockito and JPA knowledge of triplestore and/or graphstores, semantic technologies, RDF and SPARQL knowledge of XML Database (MarkLogic), XQuery experience of bug tracking, distributed version control systems and continuous integration a good understanding of software development methodologies&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Nature Publishing Group (NPG) is a division of Macmillan Publishers Ltd, dedicated to serving the academic, professional scientific and&lt;br /&gt;
medical communities. This role is in a great location in King&#039;s Cross with a lovely working environment, option for macbook, free canteen,&lt;br /&gt;
exciting and challenging work, enthusiastic colleagues, and a sensible work/life balance. The salary is excellent and negotiable for the&lt;br /&gt;
right person.&lt;br /&gt;
&lt;br /&gt;
contact marco.neumann@gmail.com&lt;br /&gt;
&lt;br /&gt;
==Sales Executives at Neo Technology, Germany ==&lt;br /&gt;
&lt;br /&gt;
Location&lt;br /&gt;
Germany (Munich)&lt;br /&gt;
&lt;br /&gt;
Overview&lt;br /&gt;
As a Sales Executive you are a hunter and closer with experience selling to developers, architects, IT Directors, and consultants within F100 firms, government agencies, and to start-ups depending on your experience. Creative, energetic and a self-starter, you understand the sales process, how to sell innovation and disruption through customer vision expansion. You can drive deals forward and compress decision cycles. You also love understanding a product in depth and then communicating that product to the marketplace. If you’re a high achiever, we want to talk to you.&lt;br /&gt;
&lt;br /&gt;
We’re looking for a track record of proven sales success with 3 -15 years of experience in software sales.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements, Qualifications, Skills &amp;amp; Abilities&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
*Qualify inbound leads representing key stakeholders, CIOs, CTO’s, IT Director’s, IT architect’s and consultant’s from various organizations&lt;br /&gt;
*Demonstrating Neo4j’s capabilities over the Web and phone&lt;br /&gt;
*Meeting/exceeding activity, pipeline, and booking targets&lt;br /&gt;
*Discipline in documenting various details including use case, purchase timeframes, next steps, and forecasting&lt;br /&gt;
*Ensuring 100% satisfaction with all customers and partners&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
*B.Sc. or M.Sc. degree required&lt;br /&gt;
*Obvious passion and people skills&lt;br /&gt;
*Extensive prior customer relationships&lt;br /&gt;
*A history of consistent over-achievement&lt;br /&gt;
*Experience in opportunity creation and the challenges of building a business&lt;br /&gt;
&lt;br /&gt;
[http://www.neotechnology.com/2012/10/neo-technology-is-hiring-sales-executives-germany/ Apply Here]&lt;br /&gt;
&lt;br /&gt;
==Software Developer with Semantic Web experience (University of Colorado, Boulder, CO)==&lt;br /&gt;
The Faculty Information System team at the University of Colorado Boulder is searching for a software developer who loves data and learning. Someone who would enjoy working on the beautiful Boulder campus, minutes from outdoor activities and Downtown Boulder. If this describes you, or someone you know, please read on.&lt;br /&gt;
&lt;br /&gt;
Our team is launching VIVO CU-Boulder, a key component of CU-Boulder&#039;s presence on the next generation Web that harnesses the possibilities of Linked Data and Open Data. The FIS team is a small, collaborative group employing agile and lean software development practices to deliver service-oriented solutions in an expanding environment. We are looking for someone who enjoys working closely with others, and who would be committed to the continuous improvement of FIS and our software development process. Candidates for the position should demonstrate success in data management and web application development using current methods and technologies. Experience with VIVO, W3C Semantic Web technologies, and/or prior contributions to an open source community are a plus.&lt;br /&gt;
&lt;br /&gt;
This position will be open until filled. Applications submitted by Wednesday June 13, 2012 will be given full consideration. For more information on this position and the full benefits package, please see http://www.jobsatcu.com/applicants/Central?quickFind=68914&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Client Development Manager (Reuters Media)==&lt;br /&gt;
 &lt;br /&gt;
Our engineering team develops business-critical applications and services for the largest news agency in the world.  We provide publishers and broadcasters with news text, photos, and video over both  Internet and satellite.  As part of a small team of high-performers, you will be our principal client-facing technical expert, and will help guide development for our products.  The ideal candidate has a deep understanding of publisher workflows, is a talented programmer, and has experience with pre- and post-sales support. &lt;br /&gt;
 &lt;br /&gt;
As our client-facing technical expert, you understand and can evangelize technical standards such as NewsML-G2 and RightsML, and keep abreast of the latest trends in publishing.  You can work with our metadata experts and provide them with our client’s perspective.  You also have an opportunity to interact with IPTC leadership in the development of new standards such as rNews.  Participation in New York City industry groups such as the NY Tech Meetup, Hacks and Hackers, and the New York Semantic Web group are all a big plus.&lt;br /&gt;
 &lt;br /&gt;
Responsibilities:&lt;br /&gt;
* Advise and assist clients on making the best use of our platform and data, including occasional visits to client sites; participate in pre-sales meetings&lt;br /&gt;
&lt;br /&gt;
* Manage technical aspects of client onboarding&lt;br /&gt;
&lt;br /&gt;
* Generate and prototype ideas for new products and features&lt;br /&gt;
&lt;br /&gt;
* Manage small development teams for new initiatives&lt;br /&gt;
&lt;br /&gt;
* Work with product managers to produce development roadmaps&lt;br /&gt;
&lt;br /&gt;
* Coordinate with development teams in Beijing&lt;br /&gt;
&lt;br /&gt;
* Work with software teams to understand client workflows and educate design and implementation of features and advise on user experience.&lt;br /&gt;
&lt;br /&gt;
Apply here&amp;lt;br&amp;gt;&lt;br /&gt;
https://toc.taleo.net/careersection/2/jobdetail.ftl?lang=en&amp;amp;job=TEC00022861&lt;br /&gt;
&lt;br /&gt;
==SEMANTIC WEB PROGRAMMER/DEVELOPER at Brown University Library in Providence, RI == &lt;br /&gt;
Brown University is in the process of implementing VIVO, an open source semantic web application that supports and facilitates research discovery within and among institutions. The Brown University Library is seeking a Semantic Web Programmer/Developer to play a vital role in the launch and support of this new campus enterprise system. This full time, permanent position is an exciting opportunity for a programmer with experience in semantic web technologies to advance a large-scale linked open data project. &lt;br /&gt;
&lt;br /&gt;
The Semantic Web Programmer/Developer is responsible for initial data ingest planning and execution, for configuration of local extensions to the application ontology and for ongoing maintenance of ontology and data. The position develops and documents scripts using XML and semantic web technologies to process data and metadata from institutional databases of record, online databases of publications and research grant information, and other sources as identified by campus stakeholders. S/He writes programs and web services to return integrated and enhanced data to institutional stakeholders in RDF, XML, JSON, and other formats for reporting analysis, archiving, and display. The position participates actively in the VIVO development network and represents Brown in the national VIVO community. The position reports to the Head, Integrated Technology Services in the Brown University Library. &lt;br /&gt;
&lt;br /&gt;
Qualifications: &lt;br /&gt;
* Bachelor’s Degree in Computer Science or Advanced Degree in Information Science; plus three to five years relevant experience &lt;br /&gt;
* Experience working with RDF data model and semantic web design principles &lt;br /&gt;
* Experience with formal ontology languages such as OWL and RDFS &lt;br /&gt;
* Experience with languages for querying RDF (e.g., SPARQL, SeRQL) &lt;br /&gt;
* Experience with one or more metadata manipulation and scripting languages: XSLT, Java, Perl, Python, or PHP. &lt;br /&gt;
* Knowledge of data extraction, mining, harvesting techniques and tools &lt;br /&gt;
* Excellent interpersonal, oral and written communication skills &lt;br /&gt;
* Ability to work independently and as a member of a team &lt;br /&gt;
&lt;br /&gt;
Preferred Qualifications: &lt;br /&gt;
Experience with Jena or other semantic web libraries &lt;br /&gt;
Experience with deployment of RDF triple stores in a production environment &lt;br /&gt;
Experience with metadata issues related to the discovery of academic resources &lt;br /&gt;
&lt;br /&gt;
To apply for this position (JOB #B01403), please visit Brown’s Online Employment website (https://careers.brown.edu), complete an application online, attach documents, and submit for immediate consideration. Documents should include cover letter, resume, and the names and e-mail addresses of three references. Review of applications will continue until the position is filled. Brown University is an Equal Opportunity/ Affirmative Action Employer. &lt;br /&gt;
&lt;br /&gt;
You can find the online description[http://library.brown.edu/about/employment.php#semantic here]&lt;br /&gt;
&lt;br /&gt;
== Linked Data Engineer at Financial Institution in NYC==&lt;br /&gt;
&lt;br /&gt;
Date:2/29/2012&lt;br /&gt;
&lt;br /&gt;
Our client, a leading financial institution, is looking for a Linked Data Engineer to work on high-profile projects within their engineering and architecture group. This position is a mission-critical opportunity, which will improve user experience globally for the client, working with linked data, RDF/OWL modeling, and Jena/Protege.&lt;br /&gt;
&lt;br /&gt;
On a daily basis, this individual will have opportunity to be involved in analysis, design, and integration of a firm-wide linked open data solution. This individual will be the on-site expert working on cutting edge technologies. This person will have the opportunity to be involved with high profile projects before they are released to the rest of the bank.&lt;br /&gt;
&lt;br /&gt;
To be qualified for this position, this individual should have strong knowledge of RDF/OWL modeling and either Jena, Protégé, or D2RQ. It would be a plus if this individual has exposure to J2EE technologies such as Tomcat, Spring, Eclipse, Maven, RMI, Java Server Faces, and JMS.&lt;br /&gt;
&lt;br /&gt;
This position is long-term with the opportunity for growth. This opportunity offers the opportunity to learn new skills and be part of a cutting edge team. Please reply if you are available for an interview within 72 hours.&lt;br /&gt;
&lt;br /&gt;
Online description: [http://information-technology.thingamajob.com/jobs/New-York/Linked-Data-Engineer/2496518 thingamajob]&lt;br /&gt;
&lt;br /&gt;
Please contact Lauren Freidhof for further information: lfreidho@teksystems.com&lt;br /&gt;
&lt;br /&gt;
== Research Engineer - The New York Times Company==&lt;br /&gt;
&lt;br /&gt;
The New York Times Company, a leading media company with 2010 revenues of $2.4 billion, includes The New York Times, the International Herald Tribune, The Boston Globe, 15 other daily newspapers and more than 50 Web sites, including NYTimes.com, Boston.com and About.com. The Company’s core purpose is to enhance society by creating, collecting and distributing high-quality news, information and entertainment.&lt;br /&gt;
&lt;br /&gt;
The New York Times Company&#039;s Research and Development group is looking for a talented developer to serve as the Research Engineer for Knowledge Management.  The New York Times has one of the world&#039;s great archives and our ideal candidate will be passionate about creating innovative prototypes, systems, standards and products from this amazing resource.  We are looking for someone with an innate curiosity and a passion for innovation and who has the ability to channel this passion into both individual and team projects.&lt;br /&gt;
 &lt;br /&gt;
Responsibilities&lt;br /&gt;
*Conceptualize and develop prototypes for innovative knowledge management products and services with an emphasis on linked data technologies&lt;br /&gt;
*Solidify existing prototypes into production products and systems&lt;br /&gt;
*Keep abreast of the evolving knowledge management landscape including the emergence of new standards, technologies and systems&lt;br /&gt;
*Deeply understand the New York Times Company&#039;s data assets and how these assets can be leveraged to create useful systems and compelling new experiences&lt;br /&gt;
 &lt;br /&gt;
Qualifications&lt;br /&gt;
*Bachelor&#039;s degree in Computer Science or related subject, Master&#039;s degree a plus&lt;br /&gt;
*3-5 years of experience in architecting, developing and delivering complex knowledge management solutions&lt;br /&gt;
*Passion for knowledge management and big data problems&lt;br /&gt;
*Flexibility to work with a variety of technologies&lt;br /&gt;
*Strong coding skills in at least one statically typed language - Java / C# / Objective C / C++&lt;br /&gt;
*Proficient coding skills in at least one dynamically typed language - Python / Ruby / etc.&lt;br /&gt;
*Experience with PHP strongly desired&lt;br /&gt;
*Proficient in at least one relational database technology - MySQL / Oracle / Access&lt;br /&gt;
*Experience or interest in learning in NoSQL database technologies - MongoDB / RDF Triplestores / etc.&lt;br /&gt;
*Familiarity with basic approaches to natural language processing&lt;br /&gt;
*Interest in the development and promotion of industry-wide digital standards&lt;br /&gt;
*Strong communication skills and experience with explaining deeply technical concepts to non-technical audiences&lt;br /&gt;
*Ability to thrive in a self-directed environment&lt;br /&gt;
 &lt;br /&gt;
The New York Times Company is an equal employment opportunity employer, and does not discriminate on the basis of race, color, religion, gender, sexual orientation, marital status, age, disability, national origin, citizenship or any other protected characteristic. The New York Times Company is committed to diversity in its most inclusive sense.&lt;br /&gt;
&lt;br /&gt;
[http://developer.nytimes.com/docs/The_Semantic_API apply here]&lt;br /&gt;
&lt;br /&gt;
==Metropolitan Musem of Art hiring a Semantic Web Developer==&lt;br /&gt;
&lt;br /&gt;
Media Lab&amp;lt;br&amp;gt;&lt;br /&gt;
Digital Media Department&amp;lt;br&amp;gt;&lt;br /&gt;
Metropolitan Museum of Art&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
the Metropolitan Museum of Art&#039;s Digital Media Department is hiring for an Information Systems Developer.&lt;br /&gt;
This position will be involved in advanced data architecture solutions, to support a variety of web and in-gallery technology.&lt;br /&gt;
&lt;br /&gt;
This work may entail:&lt;br /&gt;
- Setting up and administering triple stores, NoSQL dbs, and CMSs like Drupal&lt;br /&gt;
- designing interfaces, modules, and workflows for same&lt;br /&gt;
- Implementing collective intelligence algorithms, &lt;br /&gt;
- experimenting with new technologies, developing prototypes and proofs-of-concept&lt;br /&gt;
- and (to be honest) some drudgery, like data delivery, ETL, and report generation&lt;br /&gt;
&lt;br /&gt;
See the application on linkedin [http://www.linkedin.com/jobs?viewJob=&amp;amp;jobId=2157751&amp;amp;srchIndex=0&amp;amp;trk=njsrch_hits&amp;amp;goback=%2Efjs_information+systems+developer_*1_*1_I_us_*1_*1_1_R_true_*2_*2_*2_*2_*2_*2_*2_*2   here].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
I know many of you do more than just SemWeb work, and many of you are on this list because you like to find new ways to tackle vexing problems. That&#039;s what we&#039;re looking for.&lt;br /&gt;
&lt;br /&gt;
If you choose to submit a resume, please send it to the email address provided, but also cc: don.undeen@metmuseum.org&lt;br /&gt;
&lt;br /&gt;
== Job/Contracting Work - NYC ==&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 5/14/2011&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Hello,&lt;br /&gt;
 &lt;br /&gt;
I am looking to speak to people with Natural Language Processing/Machine Learning experience regarding a potential project to help analyze Social Media data/perform sentiment analysis.  Ideal candidate will have an academic background in different NLP/Machine Learning theories, frameworks and real-world implementations.  A background in programming is a plus.  Pay is competitive.  Please send me a message if you are interested in having a discussion to learn more about the opportunity.  Please contact me at: igorgonta@yahoo.com.&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Gluejar is Hiring ==&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039;3/1/2011&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
We have funding for 4 positions, ranging from Engineering to Marketing, in a metadata-rich problem space.&lt;br /&gt;
&lt;br /&gt;
For details, please see [http://go-to-hellman.blogspot.com/2011/03/gluejar-is-hiring.html this post on the Go To Hellman blog].&lt;br /&gt;
&lt;br /&gt;
== Metadata Manager at the Associated Press (New York City) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039;2/25/2011&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Associated Press is seeking a Metadata Manager for its New York City and/or Cranbury, NJ office locations.&lt;br /&gt;
&lt;br /&gt;
The Metadata Manager will be responsible for developing and maintaining metadata standards across AP’s diverse set of content and products. Responsibilities will include all aspects of working with metadata schema, such as information modeling, XML validation, transformation and testing. Additional duties will include maintaining schema artifacts (like XML schema definitions, metadata mapping tables and editorial workflow documentation), performing metadata and content analysis, and developing technical specifications for populating and transforming metadata structures and values.  This position reports to the Deputy Director of Schema Standards and is within the AP’s Information Management group.&lt;br /&gt;
&lt;br /&gt;
Primary responsibilities&lt;br /&gt;
&lt;br /&gt;
* Collaborate with journalists, technologists, product development and members of the Information Management team to define and document metadata schema across media types and products.&lt;br /&gt;
&lt;br /&gt;
* Gather business requirements for content and content metadata, and capture and maintain business rules related to metadata modeling and transforms.&lt;br /&gt;
&lt;br /&gt;
* Design, develop and execute content and metadata analyses, interpret results and provide written summaries and reports.&lt;br /&gt;
&lt;br /&gt;
* Work with Development and QA to specify and test schemas and transforms and to communicate changes to stakeholders.&lt;br /&gt;
&lt;br /&gt;
Knowledge, Skills and Abilities&lt;br /&gt;
&lt;br /&gt;
* Familiarity with XML and XML schema languages. Knowledge of XPath, XSLT or XQuery a plus.&lt;br /&gt;
&lt;br /&gt;
* Experience creating and implementing metadata schemas and taxonomies or controlled vocabularies.&lt;br /&gt;
&lt;br /&gt;
* Experience providing requirements to technical teams. Hands on programming experience a plus.&lt;br /&gt;
&lt;br /&gt;
* Familiarity with databases and query languages. Statistical data or content analysis experience a plus.&lt;br /&gt;
&lt;br /&gt;
* Familiarity with the publishing, entertainment or media industries a plus.&lt;br /&gt;
&lt;br /&gt;
* Excellent written and oral communications skills.&lt;br /&gt;
&lt;br /&gt;
* Demonstrated ability to work effectively across groups to achieve objectives.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039;&lt;br /&gt;
To learn more about the position and to apply:&lt;br /&gt;
https://careers.ap.org/viewjob.html?optlink-view=view-19563&amp;amp;ERFormID=newjoblist&amp;amp;ERFormCode=any&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==University of Minnesota Libraries Seeks Drupal Developers==&lt;br /&gt;
Date: 11/2/2010&lt;br /&gt;
&lt;br /&gt;
The University of Minnesota Libraries seeks two or more talented Drupal&lt;br /&gt;
software developers, for either one or two year appointments, to design and&lt;br /&gt;
support new, innovative web-based library services, systems, and tools which&lt;br /&gt;
address as well as anticipate the evolving needs of library users.&lt;br /&gt;
&lt;br /&gt;
The University Libraries are supporting multiple projects using the Drupal&lt;br /&gt;
platform.  Responsibilities could include two or more of the following areas&lt;br /&gt;
of Drupal development:&lt;br /&gt;
&lt;br /&gt;
*In collaboration with our partners in the American Indian Studies department, provide primary development support for the forthcoming Online Ojibwe Dictionary. Responsibilities include addressing both content provider and user needs in developing a robust web application using the Drupal framework.&lt;br /&gt;
*Provide development support for the University Libraries&#039; UMedia Archive (umedia.lib.umn.edu), a digital library application that provides users with access to many of the Libraries rich media collections as well as allowing for user submitted uploads. Using Drupal, work to integrate new features and support current mechanisms that help further the enhance the user experience.&lt;br /&gt;
*Assist in implementing the Drupal CMS for the main public facing web site of the University Libraries (www.lib.umn.edu) , creating customization and personalization options for library users, helping in the creation of mobile version of library web site(s), designing new sites, and using new web services technologies to improve the user experience in discovering, searching, finding, or acquiring library materials and content. Projects may also likely include further integration of library resources into the Moodle course management tool, implementation of Shibboleth identity management system, and creatively using various API&#039;s made available by Google, OCLC, Amazon, Ex Libris and other library vendors.&lt;br /&gt;
&lt;br /&gt;
For more information and to apply:&lt;br /&gt;
&lt;br /&gt;
http://employment.umn.edu/applicants/Central?quickFind=90989&lt;br /&gt;
&lt;br /&gt;
== Early Stage, Funded Start Up with $2B Pilot Customer Seeks Lead Engineer/ VP Development/Tech Assassin - CovetedList.com ==&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039;8/8/10&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Company:          CovetedList.com &lt;br /&gt;
&lt;br /&gt;
Elevator Pitch:   We provide smart data personalization services when shopping, browsing and sharing on the internet. [Translation: Our goal is to be the next generation of Nielsens or NPD Group using a combination of linked data technologies, rule based information retrieval with natural language processing and learning systems/ machine learning]&lt;br /&gt;
&lt;br /&gt;
i.e. we have a REALLY cool idea and a very disruptive business model. &lt;br /&gt;
&lt;br /&gt;
411:              Early stage, *funded* start up (with a Fortune 1000 pilot customer), looking for the Head of Development who can build a team and a platform and has a good understanding of natural language processing, data extraction, information retrieval, machine learning and familiar with semantic technologies, of course.  The ideal candidate has a MS or PhD or PhD candidate, is an out of the box academic thinker who knows how to code in Python/C++.&lt;br /&gt;
&lt;br /&gt;
We are planning to build our front end using modern, web-facing technologies (JavaScript, HTML5, etc.), and implement our back-end services and algorithms in the most dynamic programming environment that is up to the task (but momentum as of late is *strongly* leaning towards Python). The candidate will solidify this decision for us. (Note: our core libaries and ontologies are pretty built out, we are building apps for them to become living entities- which would be part of your job)&lt;br /&gt;
&lt;br /&gt;
Candidates should have experience in the product management cycle, on-time delivery, and great communication skills. Agile and XP development methodologies are good skills. Scrum is cool. If you don&#039;t know something, it is important to admit it and be psyched to learn.&lt;br /&gt;
&lt;br /&gt;
MUST BE LOCATED IN OR AROUND THE NY METRO AREA.  WE ARE IN NYC.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; charlotte@covetedlist.com&lt;br /&gt;
&lt;br /&gt;
== User Interface Developer - OrangeDog ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 5/6/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
I am putting together a &amp;quot;semantic applications&amp;quot; startup, with several semantic applications in mind.&lt;br /&gt;
&lt;br /&gt;
The first area I am looking at is data integration and user interfaces using ontologies. I have already written software for both the data integration and the user interface side of things. The software works quite nicely -- it is well beyond a prototype (it&#039;s stable, modular, not too slow, etc). The integration side of the software is much more mature than the user interface side (I am no user interface designer!). &lt;br /&gt;
&lt;br /&gt;
I have also written a detailed business plan. I am now actively looking for startup funding (the usual angel, VC, etc stuff), and I am looking for people who may be interested in joining the startup, as both employees, and as equity holders, founders, etc.  &lt;br /&gt;
&lt;br /&gt;
I am specifically interested in finding someone with interest and expertise in user interfaces. &lt;br /&gt;
&lt;br /&gt;
This isn&#039;t a user experience position, it&#039;s an engineering/coding position. The UI person will have two roles. The first is the standard one of writing UIs for the various applications we have. Pretty straightforward stuff. The second role is to build so called &amp;quot;semantic&amp;quot; user interfaces. These are interfaces that are generated from ontologies. This role is very open ended. There is basically an endless scope here for a creative user interface developer, as there are a lot of interesting things you can do when generating UIs from ontologies.&lt;br /&gt;
&lt;br /&gt;
Essentially the person I am looking for needs to be a really good UI coder, plus able to understand what ontologies bring to the table (you don&#039;t necessarily need to know this now, but be capable of learning it), and hence how users may want to interact with them visually, plus have a good visual sense, since the UI person will be in charge of the user interface side of our product suite. They must have a passion for understanding and learning how end users want to use both user interfaces and ontologies, and hence how they can support those end users -- we are relentlessly client focussed.&lt;br /&gt;
&lt;br /&gt;
To be clear -- I cannot pay anyone at this point as I haven&#039;t got funding yet. So for the moment this all has to be for equity consideration. If we do get funded this turns into a full time job.&lt;br /&gt;
&lt;br /&gt;
We are located on the west coast of the US in Los Angeles, but you don&#039;t have to be. Getting quality people is more important to us than having them in the &amp;quot;right&amp;quot; location. If you are interested, or know anyone interested, please drop me a line at graham@orangedogconsulting.com.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Web Developer - Reflexions Data ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: White Plains (Westchester/NYC)&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 4/2/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Reflexions Data is a growing company of creative innovators that develops web applications for clients in a variety of industries including publishing, marketing, retail e-commerce, internet startups and non-profit organizations. We are an equal opportunity employer (M/F/D/V).&lt;br /&gt;
&lt;br /&gt;
We&#039;ve been around for 10+ years and our office is a casual, fast-paced environment.  We offer competitive compensation and benefits including group health insurance and a 401(k) retirement plan.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
* Must be willing to work in a closely knit team environment and must demonstrate a passion for solving business problems with technology.&lt;br /&gt;
* BA/BS or MS in Computer Science or related technical discipline.&lt;br /&gt;
* Deep understanding of computer science fundamentals, including data structures, algorithms, and software design principles.&lt;br /&gt;
* Familiarity with MVC-style web development frameworks.&lt;br /&gt;
* At least 2 years of web/software application design and development experience.&lt;br /&gt;
* Extensive knowledge and experience with UNIX/Linux.&lt;br /&gt;
* Intimate familiarity with web standards and front-end technologies including XHTML, CSS, and Javascript/AJAX.&lt;br /&gt;
&lt;br /&gt;
Do you read Slashdot every day? Do you see regular expressions in your dreams or write Python code for fun? Then we encourage you to introduce yourself! This is a great mid-level position with opportunities for advancement. &lt;br /&gt;
&lt;br /&gt;
Learn more and apply here:&lt;br /&gt;
http://www.reflexionsdata.com/company/employment&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Metadata Librarian/Analyst - Bloomberg ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: Skillman, NJ&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 4/1/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Company&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Bloomberg is the world’s most trusted source of information for businesses and professionals. Bloomberg combines innovative technology with unmatched analytic, data, news, display and distribution capabilities, to deliver critical information via the BLOOMBERG PROFESSIONAL® service and multimedia platforms. Bloomberg&#039;s media services cover the world with more than 2,200 news and multimedia professionals at 146 bureaus in 72 countries. The BLOOMBERG TELEVISION® 24-hour network delivers smart television to more than 240 million homes. BLOOMBERG RADIO® services broadcast via SIRIUS XM Radio and 1worldspaceTM satellite radio globally and on WBBR 1130AM in New York. The award-winning monthly BLOOMBERG MARKETS® magazine, Bloomberg BusinessWeek magazine and the BLOOMBERG.COM® financial news and information Web site provide news and insight to businesses and investors.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Role&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Bloomberg is looking for a highly motivated individual to join a newly formed information-retrieval group.  As a Librarian/Analyst within this group, you will be part of a team dedicated to helping provide the next generation in Bloomberg Terminal usability to our clients by standardizing, organizing, and ensuring accuracy of keyword databases to support a new information-retrieval system. The successful candidate will apply taxonomy and ontology principles to create and manage metadata to improve ‘findability’ of resources, diagnose and fix inconsistencies, solve reference issues, analyze linguistic usage patterns and perform basic data analysis tasks.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;UL&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt; Bachelor&#039;s degree or equivalent work experience.  Masters in Library Science (MLS, MLIS), Bachelor of Science in Computer Science, or related technical discipline is preferred. &lt;br /&gt;
&amp;lt;li&amp;gt; 3 or more years of data analysis experience.&lt;br /&gt;
&amp;lt;li&amp;gt; Experience in use of relational databases, such as Access, MySQL, or Oracle. &lt;br /&gt;
&amp;lt;li&amp;gt; Knowledge of metadata standards such MARC, Dublin Core, etc., and the proven ability to apply metadata to large scale collections of data.&lt;br /&gt;
&amp;lt;li&amp;gt; Demonstrated ability to research and analyze problems and develop solutions.&lt;br /&gt;
&amp;lt;li&amp;gt; Strong analytic and organization skills.&lt;br /&gt;
&amp;lt;li&amp;gt; Ability to exercise independent judgment.&lt;br /&gt;
&amp;lt;li&amp;gt; Experience handling financial, economic and discrete content is helpful, but not required.&lt;br /&gt;
&amp;lt;li&amp;gt; Data modeling experience is a plus.&lt;br /&gt;
&amp;lt;/UL&amp;gt;&lt;br /&gt;
Bloomberg is an equal opportunity/affirmative action employer and we welcome applications from all backgrounds regardless of race, color, religion, sex, national origin, ancestry, age, marital status, sexual orientation, gender identity, veteran status, disability, or any other classification protected by law.&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
Please apply online at http://careers.bloomberg.com/hire/jobs/job25667.html&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Senior Software Engineer - daylife==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: New York City&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 3/18/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Come to Daylife and help us build the future of online news.  As a Senior Software Engineer, you&#039;ll develop scalable systems for organizing and delivering information to people and organizations around the world.  Add new features to our customer APIs; improve the performance of our search systems; enhance the quality of our information extraction; evolve our engineering tools to tighten our release cycles.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
* Strong problem solving spirit.&lt;br /&gt;
* Expert in (C || Java) &amp;amp;&amp;amp; Python.&lt;br /&gt;
* Expert UNIX network programming skills -- IP protocols, sockets, IPC, event-driven programming and frameworks (libevent, NIO).&lt;br /&gt;
* Expert knowledge of Linux programming and POSIX operating system concepts, and debugging in a Linux environment -- processes, pthreads, memory model, filesystems, system introspection and debugging tools.&lt;br /&gt;
* Strong software design skills.  Write well-organized, maintainable and testable code.&lt;br /&gt;
* Solid knowledge of version control with subversion; git knowledge a plus.&lt;br /&gt;
* RHEL / CentOS administration or operating experience a plus.&lt;br /&gt;
* Strong leadership, teamwork, and communication skills.&lt;br /&gt;
* Strong coaching skills; proactively shares knowledge.&lt;br /&gt;
&lt;br /&gt;
Please contact with your CV Ken Ellis (ken [at] daylife.com)&lt;br /&gt;
&lt;br /&gt;
http://www.daylife.com/&lt;br /&gt;
&lt;br /&gt;
== Senior Semantic and Search Architect - Financial Services==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location&#039;&#039;&#039;: New York City&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 3/15/2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Responsible for architecture of a content mining and semantic based product offerings for the Knowledge Management Practice Area. Researches, analyzes, and recommends technologies to develop new and enhance existing systems to support&lt;br /&gt;
semantic requirements. Serves as a technical expert and is responsible for resolving complex problems and guiding development to improve the application of search and semantic technologies.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Required Skills and Experience:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
BS in Computer Science or a related field, MS preferred&amp;lt;br&amp;gt;&lt;br /&gt;
Extensive experience with distributed software development&amp;lt;br&amp;gt;&lt;br /&gt;
Data mining and predictive modeling skills&amp;lt;br&amp;gt;&lt;br /&gt;
Strong OO knowledge&amp;lt;br&amp;gt;&lt;br /&gt;
Expertise in C#.Net or Java&amp;lt;br&amp;gt;&lt;br /&gt;
Comfortable with distributed development environments and tools&amp;lt;br&amp;gt;&lt;br /&gt;
Minimum of 10 years development experience&amp;lt;br&amp;gt;&lt;br /&gt;
Excellent communication skills&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with Search technologies – Autonomy IDOL/ FAST/ Verity / Lucene&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with mentoring software development resources.&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with scoping the technical side of client facing products that require scale and contextual consumer experiences.&lt;br /&gt;
Experience with understanding usage flows and has experience with mining historical consumer choice data to make a service more intelligent about consumer options.&amp;lt;br&amp;gt;&lt;br /&gt;
Experience in technical architectures for search applications, including data structures and algorithms to support entity extraction, disambiguation, normalization and clustering.&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with architecture models and design patterns to support development and maintenance of taxonomies and knowledge bases. Ideal candidate has been working on ways to accomplish this in an enterprise environment using open source&lt;br /&gt;
or by using data available from multiple sources.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills Desired&#039;&#039;&#039;&amp;lt;BR&amp;gt;&lt;br /&gt;
Background in natural language processing NLP&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with cloud computing platforms&amp;lt;br&amp;gt;&lt;br /&gt;
Exposure to popular open source machine learning tools&amp;lt;br&amp;gt;&lt;br /&gt;
Data harvesting experience&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with search content classification&amp;lt;br&amp;gt;&lt;br /&gt;
Experience in the publishing industry particularly in optimization&amp;lt;br&amp;gt;&lt;br /&gt;
Experience with non-relational NoSQL document oriented data stores&amp;lt;br&amp;gt;&lt;br /&gt;
Understanding of Semantic Web nomenclature and related technologies&amp;lt;br&amp;gt;&lt;br /&gt;
Key words: semantic web, Resource Description Framework RDF, Autonomy IDOL,Verity K2, data interchange formats, RDF, XML, N3, Turtle, N-Triples, RDF Schema RDFS and the Web Ontology Language OWL, semantic web stack, URI, Ontologies, NoSQL&lt;br /&gt;
&lt;br /&gt;
Please contact with your attached CV: info@kona.llc&lt;br /&gt;
&lt;br /&gt;
== User Interface Designer - Kikin==&lt;br /&gt;
&lt;br /&gt;
Kikin is searching for a User Interface Designer to work directly with our VP of Product on the user interface design. You will also be responsible for creating working product mock-ups for our key clients and partners that demonstrate Kikin&#039;s custom product capabilities.&lt;br /&gt;
&lt;br /&gt;
Our user interface designer will help us fulfill our mission of making it SIMPLE to navigate the web. This person will be the leader when it comes to making this process much easier than it is today. You will be asked to create beautiful designs that are obvious to our users. This isn&#039;t going to be a run of the mill HTML and CSS job - we&#039;re looking for someone that is going to dig in and create an experience that makes our web platform legendary. You&#039;ll be running the UI show which will include: brainstorming design and usability options, building simple wire frames, deliver clean, elegant HTML, CSS and javascript. We have ideas for how our features might look good--you know how to design them so we get unsolicited calls about how amazing the site is.&lt;br /&gt;
&lt;br /&gt;
You get a thrill from building things people love to use. You think the best interfaces are the ones that get out of the way and let people do their work, not necessarily the ones that grab the most attention. You&#039;re constantly finding the sweet spot between beauty and usability. You know users are impatient and you don&#039;t have much time to impress them. You&#039;re oozing with creativity. When someone asks you how you&#039;d design something, you immediately think of ten completely different options. You can mock them up in minutes, not hours, and you can objectively evaluate the most usable. You&#039;re as comfortable drawing sketches on paper napkins as you are throwing layers around in Adobe Creative Suite 4. You feed off the energy and enthusiasm of others. Simply put, you love the challenge of building a great user experience from inception to delivery.&lt;br /&gt;
What you need:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* 2+ years experience with HTML/CSS development and graphic design for professional web projects&lt;br /&gt;
* Detailed knowledge of HTML/CSS standards &lt;br /&gt;
* Basic javascript proficiency &lt;br /&gt;
* Graphic Design (You really know your way around a professional graphics package)&lt;br /&gt;
* Strong understanding of web application usability&lt;br /&gt;
* Portfolio of some of your past work&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;About Us&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
At Kikin, we’re driven to revolutionize online &amp;amp; mobile navigation, services, and advertising.  Using proprietary client-side technology, kikin provides on-the-fly, browser-based enhancements that deliver personalization, increased relevance, richer merchandising and content, and integrated services-- all without users having to change any of their existing behavior.  kikin supports all popular search engines, major service providers, browsers, operating systems, etc.&lt;br /&gt;
&lt;br /&gt;
Kikin offers a fantastic environment and a team of bright, dynamic people from all over the world. We are dedicated to fundamentally changing the way people navigate and use the Internet and having a great time doing it.&lt;br /&gt;
&lt;br /&gt;
While maintaining a low profile, we’re testing in beta and continue to win key partnerships with major content, commerce, service providers, OEM distributors and developers. Kikin was founded by seasoned entrepreneurs with successful track records.  Operations have been established in the U.S. (NYC) and Europe (Berlin), and will be established in China (Shanghai) and Japan (Tokyo) by the end of Q4CY09. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
This position is full-time in New York City (SOHO).  This is not a telecommute or contract role.&lt;br /&gt;
No third-party, sub-contractors/agencies. Unfortunately sponsorship are Not available.&lt;br /&gt;
&lt;br /&gt;
We will require credible work references/background check.&lt;br /&gt;
Salary commensurate with experience.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Team Lead/ Back End Search, Data Engineer - kikin==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The Java Developer position at kikin is working on enhancing kikin’s next-generation internet application and services. &lt;br /&gt;
Software development tasks are focused on information retrieval, data mining, relationship mapping and collaborative filtering.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Responsibilities&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Lead a small team that will:&lt;br /&gt;
* Participate in all stages of development including Design, Implementation, and Testing of our next-generation Federated search and Personalization back end&lt;br /&gt;
* Integrate Content, Commerce,  or Service Partner data and functionality into our backend&lt;br /&gt;
* Collaborate with analysts and product management on engaging features and algorithms&lt;br /&gt;
* Work on specifications that address evolving business requirements, user interfaces, process flow, performance, and scalability&lt;br /&gt;
* Reports to CTO&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills required&#039;&#039;&#039;&lt;br /&gt;
* 4+ years of Advanced Java; Expertise in performance-oriented Java&lt;br /&gt;
* Service Oriented Architecture (SOA) experience is paramount&lt;br /&gt;
* SOAP, REST, Caching, Clustering, Distributed Computing, XML Object Mapping, Web Services&lt;br /&gt;
* Experience with Wicket, Spring, AJAX, DHTML, JavaScript, jQuery, CSS&lt;br /&gt;
* Expert XML, MySQL, Apache, Resin, Linux&lt;br /&gt;
* Experience with Compass, Lucene, SOLR desired&lt;br /&gt;
* High throughput production experience preferred&lt;br /&gt;
* Test-driven development and iterative self-correction must be second nature&lt;br /&gt;
* Self-directed, highly motivated, and able to work in a fast paced startup environment&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Database Architect - Edifice==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Northern New Jersey&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; February 20, 2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity:&#039;&#039;&#039; Our innovative company is looking for a brilliant database architect to design the back-end of a new SaaS product.  If this is you or someone you know, please contact me directly:  nbruce[at]EdificeInfo.com&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Senior Software Engineer, machine learning &amp;amp; distributed computing experience, Java, OO, Linux ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York City&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; January 30, 2010&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contract&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We are a consulting and recruiting firm. Our client, an innovative advertising startup, is seeking a creative Senior Software Engineer with machine learning and distributed computing experience. This is a hands on position that requires extensive development for release in a production environment. The ideal candidate will have good analytical and troubleshooting skills, fluency in coding and excellent communication skills.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Responsibilities&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Work with an open/friendly team of engineers and research scientists&lt;br /&gt;
* Contribute to product vision/direction&lt;br /&gt;
* Create robust high-transaction production applications&lt;br /&gt;
* Develop prototypes for research projects&lt;br /&gt;
* Production application development based on research&lt;br /&gt;
* Production software and environment troubleshooting&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* BS in Computer Science or a related field, MS preferred&lt;br /&gt;
* Extensive experience with distributed software development&lt;br /&gt;
* Data mining and predictive modeling skills&lt;br /&gt;
* Strong OO knowledge&lt;br /&gt;
* Expert in Java&lt;br /&gt;
* Comfortable with distributed development environments and tools&lt;br /&gt;
* Expert in the Linux/OS X command line&lt;br /&gt;
* Proven list of shipped products&lt;br /&gt;
* Solid math background&lt;br /&gt;
* Excellent communication skills&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Desired&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Background in natural language processing NLP&lt;br /&gt;
* Previous startup experience&lt;br /&gt;
* Experience with cloud computing platforms&lt;br /&gt;
* Exposure to popular open source machine learning tools&lt;br /&gt;
* Data harvesting experience&lt;br /&gt;
* Experience with search content classification or spam detection&lt;br /&gt;
* Experience in the advertising industry particularly in optimization&lt;br /&gt;
* Experience with non-relational NoSQL document oriented data stores&lt;br /&gt;
* Understanding of Semantic Web nomenclature and related technologies&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Critical Key words: &#039;&#039;&#039; semantic web, Resource Description Framework RDF, a variety of data interchange formats, RDF,XML, N3, Turtle, N-Triples, RDF Schema RDFS and the Web Ontology Language OWL,semantic html, Resource Description Framework RDF, Web Ontology Language OWL, Extensible Markup Language XML. RDF, OWL, and XML, semantic web stack, uri, rdfs, ontologies, NoSQL, Friend of a Friend or FoaF, DBpedia&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;good to have key words:&#039;&#039;&#039; trust metric, Friendster, livejournal, virtual communities, Slashdot, karma, rummble.com, transitivity, advogato, reputation systems, subjective logic, applied computational trust, trust management, trust engines, risk engines, trustworthy recommenders, trust-based collaborative filtering, reputation and recommendation, evidence gathering, security through collaboration, technical trust, user trust, network of trust, web of trust, impact of social networking on trust and security, virtual and self-organization, decentralized identity management, trust metrics analysis, aggregation analytics, marketing metrics, social web, viral web, Trust management, Trustos, Nuglets in mobile ad-hoc networks, Slashdot.org&#039;s Karma, Ebay&#039;s feedback rating, FOAF trust module, Free Haven direct and meta trust; direct trust, Advogato&#039;s trust metric, probabilistic trust, System trust, interpersonal and self trust&lt;br /&gt;
&lt;br /&gt;
Please send resume to molly [at] fremontconsulting.com&lt;br /&gt;
&lt;br /&gt;
for additional opportunities: visit us on Facebook:&lt;br /&gt;
&lt;br /&gt;
http://www.facebook.com/business/dashboard/#/pages/Elk-Grove-CA/Fremont-Consulting/321998995452&lt;br /&gt;
&lt;br /&gt;
* Compensation: hourly&lt;br /&gt;
* This is a contract job.&lt;br /&gt;
* OK for recruiters to contact this job poster.&lt;br /&gt;
* Please, no phone calls about this job!&lt;br /&gt;
* Please do not contact job poster about other services, products or commercial interests.&lt;br /&gt;
&lt;br /&gt;
== Python Software Engineer  ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Company&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
TLists builds search and content management tools that enable leading media companies and the mass-market to take full advantage of Twitter Lists as a brand building content distribution channel.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We&#039;re looking for an intensely talented python engineer to work with our CTO (Stanford Ph.D.) and development crew to contribute to our search platform and API. We like candidates with broad skills, but to stand out you should have an exceptional record at solving hard analytic problems, and building innovative web or data mining applications.&lt;br /&gt;
&lt;br /&gt;
The ideal engineer:&lt;br /&gt;
- Has a BA or MS in computer science or a related field (E.g., cognitive science, math, statistics, or linguistics, etc.)&lt;br /&gt;
&lt;br /&gt;
- Is experienced with python, and using python for web engineering projects (e.g., search, data aggregation, data APIs, python integration with lucene/hadoop, systems architecture, AWS, etc)&lt;br /&gt;
&lt;br /&gt;
- Has interest and/or experience in some area of data mining (e.g., machine learning, computational linguistics, statistics, etc.)&lt;br /&gt;
&lt;br /&gt;
- Is based in the New York City Area, or willing to relocate&lt;br /&gt;
&lt;br /&gt;
The position comes with equity, competitive salary, health benefits, a great comp&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact Information:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Please contact marguerite (at) tlists (dot) com to learn more about this opportunity.&lt;br /&gt;
&lt;br /&gt;
== Python Web Developer  ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Location:&#039;&#039;&#039; New York, NY USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Company&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
TLists builds search and content management tools that enable leading media companies and the mass-market to take full advantage of Twitter Lists as a brand building content distribution channel.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We&#039;re looking for an intensely talented web developer to join our CTO (Stanford Ph.D.) and development crew in New York City.&lt;br /&gt;
&lt;br /&gt;
We like candidates with broad skills, but to stand out you should have strong coding skills, a keen aesthetic sense, and an exceptional record at building innovative and beautiful web sites.&lt;br /&gt;
&lt;br /&gt;
Desirable experience includes:&lt;br /&gt;
- BA or MS in a technical field (E.g., computer science, math, cognitive science, etc.)&lt;br /&gt;
&lt;br /&gt;
- Python (or similar interpreted languages)&lt;br /&gt;
&lt;br /&gt;
- Django (or similar MVC frameworks)&lt;br /&gt;
&lt;br /&gt;
- Javascript and CSS coding (e.g., jquery, YUI, prototype, etc.)&lt;br /&gt;
&lt;br /&gt;
The position comes with equity, competitive salary, health benefits, a great computing setup, and flexible working conditions.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact Information&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Please contact marguerite (at) tlists (dot) com to learn more about this opportunity.&lt;br /&gt;
&lt;br /&gt;
==  VP of Engineering / Lead Python Software Engineer ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our Company&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
TLists builds search and content management tools that enable leading media companies and the mass-market to take full advantage of Twitter Lists as a brand building content distribution channel.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The Opportunity&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We&#039;re looking for a talented and experienced engineer to work alongside our CTO (Stanford Ph.D.) to lead the day-to-day efforts of the engineering team (search/API and web development groups).&lt;br /&gt;
&lt;br /&gt;
The ideal engineer:&lt;br /&gt;
- Has an M.S. or Ph.D. degree in CS or a related technical field (E.g., cognitive science, math, statistics, linguistics, etc.) and an exceptional record at solving challenging problems.&lt;br /&gt;
&lt;br /&gt;
- Has had a minimum of 5 years of professional coding experience, ideally with top search / advertising tech / or data mining oriented companies, as well as experience managing others in a team.&lt;br /&gt;
&lt;br /&gt;
- Is experienced with some or all of the following: data mining and statistics, systems architecture, search technology, building large/scalable web services, hadoop, lucene/solr, AWS, Twitter API, Twisted framework.&lt;br /&gt;
&lt;br /&gt;
- Is deeply proficient in python (Django exp a plus), and is quick to learn new languages and technologies (Java and Javascript also pluses).&lt;br /&gt;
&lt;br /&gt;
- Is based in the New York City Area, or willing to relocate.&lt;br /&gt;
&lt;br /&gt;
This position comes with significant equity, competitive salary, health benefits, a great computing setup, and flexible working conditions.&lt;br /&gt;
&lt;br /&gt;
TLists favors a relatively non-hierarchical working environment. As lead, you would spend roughly half your time developing code and the other half mentoring and coordinating the others on the team.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact Information&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Please contact marguerite (at) tlists (dot) com to learn more about this opportunity.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Senior Applications Developer-NYPL==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; October 8, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
http://jobs-nypl.icims.com/jobs/5655/job&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;General Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Under the direction of the Managing Director, NYPL Digital Labs:&lt;br /&gt;
&lt;br /&gt;
* Supports the implementation of a multi-instance Fedora repository&lt;br /&gt;
* Codes, integrates, and maintains services and applications that support digital object ingest, preservation, search, discovery, distribution, and &lt;br /&gt;
* Designs, implements, tests, and writes documentation of custom software applications&lt;br /&gt;
* Integrates and extends various open-source solutions&lt;br /&gt;
* Manipulates large metadata sets&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Eligibility Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Bachelor&#039;s degree in Computer Science or relevant field; advanced degree preferred&lt;br /&gt;
* 3-5 years of related experience.&lt;br /&gt;
* Extensive experience with relational databases, database design, and fluency in SQL&lt;br /&gt;
* Strong Java skills and object-oriented design experience, including knowledge of core libraries, servlets, JDBC Experience with Apache Web server and Tomcat application server&lt;br /&gt;
* Knowledge of RESTful architectures and HTTP; familiarity with RDF, OWL and triplestores preferred&lt;br /&gt;
* Experience with other programming languages, such as PHP, Python, or Ruby preferred&lt;br /&gt;
* Experience with Web services technologies (SOAP/WSDL) preferred&lt;br /&gt;
* Experience with the Fedora repository software is preferred&lt;br /&gt;
&lt;br /&gt;
==4 OPEN Ph.D. POSITIONS at Hasso-Plattner-Institute (HPI)==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; September 25, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Berlin, Germany&lt;br /&gt;
&lt;br /&gt;
We offer 4 OPEN Ph.D. POSITIONS at Hasso-Plattner-Institute (HPI), Potsdam (Germany) starting on the fourth quarter of 2009&lt;br /&gt;
&lt;br /&gt;
Hasso-Plattner-Institute (HPI) is a privately financed institute&lt;br /&gt;
affiliated with the University of Potsdam, Germany.&lt;br /&gt;
The Institute&#039;s founder and benefactor Professor Hasso Plattner,&lt;br /&gt;
who is also co-founder and chairman of the supervisory board of SAP AG,&lt;br /&gt;
has created an opportunity for students to experience a unique education in IT systems engineering&lt;br /&gt;
in a professional research environment with a strong practice orientation.&lt;br /&gt;
(for more information on HPI, c.f. http://www.hpi.uni-potsdam.de/ )&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Project Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
MEDIAGLOBE is part of the THESEUS research program,&lt;br /&gt;
initiated by the German Federal Ministry of Economy and Technology (BMWi),&lt;br /&gt;
with the goal of developing a new Internet-based infrastructure&lt;br /&gt;
in order to better use and utilize the knowledge available on the Internet.&lt;br /&gt;
The focus of the research program is on semantic technologies,&lt;br /&gt;
which determine contents (words, images, sounds, and videos)&lt;br /&gt;
not through conventional methods (e.g., combinations of letters)&lt;br /&gt;
but which are able to recognize and place the meaning of a content in its proper context.&lt;br /&gt;
MEDIAGLOBE deals with digitalization, analysis, and semantic retrieval&lt;br /&gt;
of historical, documentary audiovisual content.&lt;br /&gt;
(for more information on MEDIAGLOBE, c.f. http://theseus-programm.de/theseus-mittelstand-2009/ )&lt;br /&gt;
&lt;br /&gt;
The ideal candidate holds a MS degree in Computer Science or related field&lt;br /&gt;
and is able to consider both theoretical and practical/ implementation aspects in her/his work.&lt;br /&gt;
Fluent English communication and programming skills are fundamental requirements.&lt;br /&gt;
Preferably the candidate has a background in one of the following fields:&lt;br /&gt;
&lt;br /&gt;
* semantic web technologies&lt;br /&gt;
* knowledge representations and ontology engineering&lt;br /&gt;
* audiovisual retrieval and analysis&lt;br /&gt;
* semantic search&lt;br /&gt;
* innovative web development&lt;br /&gt;
* user interface design for audiovisual content&lt;br /&gt;
&lt;br /&gt;
The position starts as soon as possible and is full-time (40h/week)&lt;br /&gt;
for the duration of the project until Oct 2011.&lt;br /&gt;
Review of applications will begin immediately and will continue until the position is filled.&lt;br /&gt;
The successful candidate will work tightly with international partners&lt;br /&gt;
and has the possibility to pursue PhD work within the scope of the project.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;How to apply:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Excellent candidates are invited to apply with:&lt;br /&gt;
&lt;br /&gt;
* Curriculum vitae and copies of degree certificates/transcripts,&lt;br /&gt;
* Writing samples/copies of relevant scientific papers (e.g. thesis,etc.),&lt;br /&gt;
* Letters of recommendation.&lt;br /&gt;
&lt;br /&gt;
Please send your application in PDF format, indicating in the subject &amp;quot;Application for PhD position&amp;quot;&lt;br /&gt;
via email or traditional mail to the following contact.&lt;br /&gt;
&lt;br /&gt;
Contact and application:&amp;lt;br&amp;gt;&lt;br /&gt;
Harald Sack&amp;lt;br&amp;gt;&lt;br /&gt;
Hasso-Plattner-Institut für Softwaresystemtechnik GmbH&amp;lt;br&amp;gt;&lt;br /&gt;
Universität Potsdam&amp;lt;br&amp;gt;&lt;br /&gt;
Prof.-Dr.-Helmert-Str. 2-3&amp;lt;br&amp;gt;&lt;br /&gt;
D-14482 Potsdam, Germany&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
phone: 	+49 (0)331-5509-527&amp;lt;br&amp;gt;&lt;br /&gt;
fax: 	+49 (0)331-5509-325&amp;lt;br&amp;gt;&lt;br /&gt;
email: 	harald.sack[AT]hpi.uni-potsdam.de&amp;lt;br&amp;gt;&lt;br /&gt;
web:   	http://www.hpi.uni-potsdam.de/meinel/persons/sack.html&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Exciting, semantic web search startup looking for Web/Data Mining Engineer==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; August 28, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Menlo Park, CA, USA&lt;br /&gt;
&lt;br /&gt;
I’m reaching out to the group on behalf of a Series B semantic web search startup in Menlo Park&lt;br /&gt;
that I’m currently working with who is looking for a Sr. Web/Data Mining Engineer&lt;br /&gt;
to join their 30-person team on a full-time basis.&lt;br /&gt;
This is a priority hire for them, and they’re looking to move quickly.&lt;br /&gt;
If you’re interested and would like more details about this opportunity,&lt;br /&gt;
feel free to reply directly to me, and I’ll get you the pertinent information accordingly.&lt;br /&gt;
&lt;br /&gt;
Thanks!&lt;br /&gt;
&lt;br /&gt;
-Donald&lt;br /&gt;
&lt;br /&gt;
35095&lt;br /&gt;
&lt;br /&gt;
Donald James&lt;br /&gt;
&lt;br /&gt;
Technical Recruiter&amp;lt;br&amp;gt;&lt;br /&gt;
Tel (408) 727-9000&amp;lt;br&amp;gt;&lt;br /&gt;
Fax (408) 716-8882&amp;lt;br&amp;gt;&lt;br /&gt;
dkj@terransys.com&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
http://www.terransys.com&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Follow Terran Systems on Twitter!&amp;lt;br&amp;gt;&lt;br /&gt;
www.twitter.com/terran_systems&amp;lt;br&amp;gt;&lt;br /&gt;
Are you Linked In? Let&#039;s link...or just view my profile:&amp;lt;br&amp;gt;&lt;br /&gt;
http://www.linkedin.com/in/dkjames&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Semantic Web Developer Java/Scala - Zurich==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Zürich, Switzerland&lt;br /&gt;
&lt;br /&gt;
Trialox.org is a startup company at the University of Zurich.&lt;br /&gt;
Our aim is to produce an open source platform for semantic applications,&lt;br /&gt;
as well as a content management system tailored to the needs&lt;br /&gt;
of international not-for-profit organizations.&lt;br /&gt;
&lt;br /&gt;
To extend our developer team,&lt;br /&gt;
we are looking for a Senior Developer experienced in programming Semantic Web applications in Java or Scala.&lt;br /&gt;
You should have a solid knowledge around Semantic Web technology&lt;br /&gt;
and ideally be experienced with the following technologies and methodologies:&lt;br /&gt;
&lt;br /&gt;
* J2SE&lt;br /&gt;
* OSGi (with Declarative Services)&lt;br /&gt;
* Maven&lt;br /&gt;
* JAX-RS&lt;br /&gt;
* Scala&lt;br /&gt;
* Scrum&lt;br /&gt;
* Test-Driven Development&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Our expectation:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Apart from the technical skills,&lt;br /&gt;
we expect you to be motivated to help a small committed team to succeed.&lt;br /&gt;
This means that you know how to effectively explain your design choices,&lt;br /&gt;
as well as the occasional agreement to a compromise in order to get things done.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;What you can expect from us:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
We offer a great working environment with cutting-edge technologies,&lt;br /&gt;
in a team that&#039;s committed to open source and Semantic Web standards.&lt;br /&gt;
Our ties to the University allow a continuous exchange with the latest research.&lt;br /&gt;
As a company, we are committed to maintaining a great work-life balance.&lt;br /&gt;
We offer flexible working hours as well as insurance coverage beyond the legal requirements.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;To Apply:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
If you are an excellent Software Developer and would like to apply for this position,&lt;br /&gt;
please send your CV and a cover letter demonstrating your experience to:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;hr@trialox.org&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
or&lt;br /&gt;
&lt;br /&gt;
trialox ag&amp;lt;br&amp;gt;&lt;br /&gt;
Tsuyoshi Ito&amp;lt;br&amp;gt;&lt;br /&gt;
Binzmuehlestrasse 14&amp;lt;br&amp;gt;&lt;br /&gt;
CH-8050 Zürich&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For any questions do not hesitate to contact me.&lt;br /&gt;
&lt;br /&gt;
Regards,&amp;lt;br&amp;gt;&lt;br /&gt;
Reto Bachmann&lt;br /&gt;
&lt;br /&gt;
--&lt;br /&gt;
Reto Bachmann-Gmür&amp;lt;br&amp;gt;&lt;br /&gt;
trialox.org&amp;lt;br&amp;gt;&lt;br /&gt;
Tel: +41445005015&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Senior Developer (Convention Center, in Washington, DC)==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Washington, DC, USA&lt;br /&gt;
&lt;br /&gt;
[http://washingtondc.craigslist.org/doc/sof/1260179096.html Original Posting at Craigslist]&lt;br /&gt;
&lt;br /&gt;
Semantic technologies startup in downtown Washington, DC, seeks Senior Java Developer&lt;br /&gt;
to join engineering team working on challenging problems&lt;br /&gt;
in knowledge representation, AI, and related fields.&lt;br /&gt;
&lt;br /&gt;
Qualified applicants will have&lt;br /&gt;
&lt;br /&gt;
* BS in computer science or related field; MS or PhD, ideally;&lt;br /&gt;
&lt;br /&gt;
* 5+ years of software development experience, including demonstrable Java expertise;&lt;br /&gt;
&lt;br /&gt;
* very strong software development, engineering skills;&lt;br /&gt;
&lt;br /&gt;
* background in logic, AI, KR, automated reasoning, automated planning, or related subfields, including machine learning, statistical inference; experience with Semantic Web standards, particularly OWL, ideal;&lt;br /&gt;
&lt;br /&gt;
* good verbal &amp;amp; written communication skills;&lt;br /&gt;
&lt;br /&gt;
* authorization to work full-time in the US, at our downtown DC office (telecommuting is a possibility but only in extraordinary circumstances).&lt;br /&gt;
&lt;br /&gt;
A successful applicant must be comfortable with&lt;br /&gt;
&lt;br /&gt;
* developing cutting-edge automated reasoning or automated planning systems; see Pellet (http://clarkparsia.com/pellet) or HotPlanner (http://clarkparsia.com/planner);&lt;br /&gt;
&lt;br /&gt;
* reading CS and other literature and implementing algorithms to production-quality;&lt;br /&gt;
&lt;br /&gt;
* working on bleeding-edge R&amp;amp;D projects -- i.e., think solving &amp;quot;DARPA Hard&amp;quot; problems;&lt;br /&gt;
&lt;br /&gt;
* presenting complex ideas simply to colleagues &amp;amp; customers at conferences like SemTech, ISWC, DL Workshop, OWLED, etc.;&lt;br /&gt;
&lt;br /&gt;
* familiar with open source development culture, methodologies, and process.&lt;br /&gt;
&lt;br /&gt;
About Clark &amp;amp; Parsia&lt;br /&gt;
&lt;br /&gt;
Clark &amp;amp; Parsia LLC is a bootstrap startup focusing on solving infrastructure-level problems&lt;br /&gt;
in semantic web and related areas;&lt;br /&gt;
we focus on automated reasoning, automated planning, information integration,&lt;br /&gt;
and ontology-based information systems.&lt;br /&gt;
We do advanced R&amp;amp;D in the semantic technologies area and, increasingly,&lt;br /&gt;
are focused on commercializing our research into products.&lt;br /&gt;
&lt;br /&gt;
We work in a relaxed, comfortable environment where good communication,&lt;br /&gt;
world-class coffee &amp;amp; espresso, great benefits, and competitive salaries rule the day. (Joel Test Score: 9/12)&lt;br /&gt;
&lt;br /&gt;
Clark &amp;amp; Parsia LLC is an equal opportunity employer.&lt;br /&gt;
&lt;br /&gt;
* Location: Convention Center&lt;br /&gt;
* Compensation: Commensurate with skills &amp;amp; experience&lt;br /&gt;
* Principals only. Recruiters, please don&#039;t contact this job poster.&lt;br /&gt;
* Please, no phone calls about this job!&lt;br /&gt;
* Please do not contact job poster about other services, products, or commercial interests.&lt;br /&gt;
&lt;br /&gt;
==DO YOU HAVE A BEAUTIFUL MIND?==&lt;br /&gt;
&#039;&#039;Seeking exceptional social semantic web tech start-up&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
Do you want to disrupt social networking?&lt;br /&gt;
Do you believe that someone has yet to bring a true sense&lt;br /&gt;
of purpose, relevancy and broad-scale yet efficient utility to social networking?&lt;br /&gt;
Do you have a beautiful mind that can see what is possible and then build it?  &lt;br /&gt;
&lt;br /&gt;
I was a co-founder of GoTo.com (later Overture),&lt;br /&gt;
where we revolutionized search by inventing the paid search business model.&lt;br /&gt;
Now, after a variety of other adventures&lt;br /&gt;
including a few years working on the DOE public school reform initiative in NYC&lt;br /&gt;
and some time in the Internet space in China,&lt;br /&gt;
I am developing a concept to drive social networking to the next level,&lt;br /&gt;
moving beyond the existing paradigm of reinforcing existing networks&lt;br /&gt;
to one where the network grows in new and useful ways&lt;br /&gt;
by making meaningful introductions to people one doesn’t know.  &lt;br /&gt;
&lt;br /&gt;
I am looking for an exceptional technical co-founder&lt;br /&gt;
who will provide the technical vision to my product vision -&lt;br /&gt;
someone who has extensive experience in delivering complex web-based applications of significant scale –&lt;br /&gt;
both front end and back end -&lt;br /&gt;
and who is extremely well-versed in the web technologies and standards of the “social semantic web”&lt;br /&gt;
and is capable of judiciously applying them.&lt;br /&gt;
Someone who wants to build a disruptive product&lt;br /&gt;
through the innovative application of bleeding edge web 2.0/3.0 technology.&lt;br /&gt;
And someone who is an excellent collaborator.&lt;br /&gt;
&lt;br /&gt;
You will be responsible for crystallizing the product strategy with me,&lt;br /&gt;
determining the product architecture;&lt;br /&gt;
evaluating the right platforms, database strategy and management, languages and technologies to deploy;&lt;br /&gt;
and hiring, training and leading an ace technology team to build it.&lt;br /&gt;
You should be very proficient in JAVA, AJAX, Flash, and PHP to name just a few.&lt;br /&gt;
You should embrace interoperability and data portability&lt;br /&gt;
(and the belief that users own their data)&lt;br /&gt;
and have experience with OpenID, OAuth, XFN, FOAF, XRD, XMPP, RSS, REST,&lt;br /&gt;
the Portable Contacts protocal, and the other building blocks of data portability.&lt;br /&gt;
And lastly, but importantly, you should have hands on (applied) experience in social network analysis&lt;br /&gt;
and in using machine learning, natural language and text data mining technologies&lt;br /&gt;
and other relevant semantic web tools to solve hard problems.&lt;br /&gt;
You will have a huge white canvas to work on as you will be coming in on the ground floor of this opportunity.&lt;br /&gt;
The one criterion – be passionate about building a powerful and simple platform,&lt;br /&gt;
with an elegant and intuitive UI and consumer experience.&lt;br /&gt;
&lt;br /&gt;
The position is based in NY&lt;br /&gt;
(no long-distance applications please unless you are prepared to move yourself to NYC or travel back and forth)&lt;br /&gt;
and is eligible for possible &amp;quot;CTO&amp;quot; status for the person with the right experience.&lt;br /&gt;
Ideally you are in a position to take no (or nominal) salary in favor of significant equity.  &lt;br /&gt;
&lt;br /&gt;
If this resonates with you, let’s talk.&lt;br /&gt;
I am looking for that rare person everyone is looking for –&lt;br /&gt;
someone who brings together that powerful mix of deep technical experience,&lt;br /&gt;
a strong personal financial situation, sound business judgment&lt;br /&gt;
and the ability to work productively with others.&lt;br /&gt;
But changing the world takes amazing people and I am prepared to find that person.&lt;br /&gt;
Please send me your resume at &#039;&#039;&#039;stephanie@sarka.us&#039;&#039;&#039; with a few words about your areas of expertise.&lt;br /&gt;
  &lt;br /&gt;
p.s.  No web shops or other third party providers, please.&lt;br /&gt;
&lt;br /&gt;
== Looking for a Partner/CTO ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 05-19-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
Equipped with a business idea in the semantic web world for enterprises and a business background,&lt;br /&gt;
I am looking for a partner/CTO to take the technical lead on the startup that I am working on.&lt;br /&gt;
This is an opportunity to get involved with a pre-funding start-up as a founding partner.&lt;br /&gt;
If you are available, open for a non-paid / equity job,&lt;br /&gt;
and are interested in getting involved in a start-up adventure, please get in touch.&lt;br /&gt;
I’m happy to discuss the project in greater detail.&lt;br /&gt;
&lt;br /&gt;
Please contact me at boaz_cn@hotmail.com (Boaz)&lt;br /&gt;
&lt;br /&gt;
== Chief Architect, Financial Times Search (FTS) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; April 13, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039; Chief Architect &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Reports To:&#039;&#039;&#039; Chief Technology Officer &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Stamford, CT &lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Here’s something to take note of if you are experienced in search and intelligent systems.   &lt;br /&gt;
&lt;br /&gt;
If you are creative, AND quantitative, we have an extraordinary opportunity for you.&lt;br /&gt;
Think of us as the place for which you got all of that education and experience.&lt;br /&gt;
If you think you might just be the right person who is interested in aiding people&lt;br /&gt;
with the burning questions of their day,&lt;br /&gt;
leveraged and fortified with sophisticated computer software&lt;br /&gt;
and have a demonstrated interest in AI or natural language processing,&lt;br /&gt;
we have a place for you to expand your horizons.   &lt;br /&gt;
&lt;br /&gt;
Financial Times Search (FTS)&lt;br /&gt;
(part of the Financial Times Group in turn part of Pearson the largest education publisher in the world)&lt;br /&gt;
is a new business utilizing a unique and proprietary search platform.&lt;br /&gt;
Our new stand-alone search product is called Newssift&lt;br /&gt;
and will soon be indexing tens of thousands of sources and many millions of articles for business people.&lt;br /&gt;
Our Beta is up and working at Newssift.com.   &lt;br /&gt;
 &lt;br /&gt;
This startup stands to deliver targeted search results&lt;br /&gt;
with a level of accuracy and relevance unmatched on the web today&lt;br /&gt;
and we would like to engage an experienced technical manager&lt;br /&gt;
to help plan and scope the engineering and significant product aspects of this business.&lt;br /&gt;
Check the product and the reviews out on the web and you will see we are on to something.    &lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Position Summary:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Responsible for architecture of a semantic based search product.&lt;br /&gt;
Researches, analyzes, and recommends technologies&lt;br /&gt;
to develop new and enhance existing systems to support semantic search applications.&lt;br /&gt;
Serves as a technical expert and is responsible for resolving complex problems&lt;br /&gt;
and guiding development to improve www.newssift.com, a consumer facing search application.&lt;br /&gt;
The role, while focusing on search,&lt;br /&gt;
will also interface with building systems around a subscription based product&lt;br /&gt;
as well as monitoring and measurement tools.&lt;br /&gt;
A command of building systems around where to go to answer a business person’s questions&lt;br /&gt;
is the main focus of the job.   &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Required Skills and Experience:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* Experience with mentoring software development resources.&lt;br /&gt;
* Experience with scoping the technical side of consumer facing products that require scale and contextual consumer experiences. &lt;br /&gt;
* Experience with understanding traffic flows and consumer choices.   Ideally the candidate has experience with mining historical consumer choice data to make a service more intelligent about consumer options. &lt;br /&gt;
* Experience in technical architectures for search applications, including data structures and algorithms to support entity extraction, disambiguation, normalization and clustering.&lt;br /&gt;
* Experience with architecture models and design patterns to support development and maintenance of taxonomies and knowledge bases.  Ideal candidate has been working on ways to do this in an open wiki way or by using data available from multiple sources. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Required Minimum Education:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* Master’s Degree in Computer Sciences or computational linguistics or equivalent field or related experience. &lt;br /&gt;
&lt;br /&gt;
Think about it.&lt;br /&gt;
One of the world’s largest companies, committed to education and information,&lt;br /&gt;
is on the leading edge of introducing a meaning based search and query tool.&lt;br /&gt;
Indeed we have a platform off of which to work, and a beachhead in the market.&lt;br /&gt;
If you have your sights set high – you should follow up with this one.   &lt;br /&gt;
&lt;br /&gt;
Only highly motivated and very smart folks need apply to Susan Blank at susankb48@msn.com&lt;br /&gt;
&lt;br /&gt;
==Web Tier/UI Developer (AJAX and JavaScript), Financial Times Search (FTS) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; April 13, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;          Web Tier/UI Developer (AJAX and JavaScript) &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Reports To:&#039;&#039;&#039;                 Senior Director Software Development &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039;                     Stamford, CT&lt;br /&gt;
&lt;br /&gt;
Financial Times Search (FTS) is a new business utilizing a unique and proprietary search engine&lt;br /&gt;
consistent with the heritage of the Financial Times.&lt;br /&gt;
FTS yields targeted search results with a level of accuracy and relevance unmatched on the web today.&lt;br /&gt;
FTS is a member of the Financial Times Group, part of Pearson PLC,&lt;br /&gt;
the world’s largest educational publisher&lt;br /&gt;
and owner of familiar businesses, including Prentice Hall and Penguin Books. &lt;br /&gt;
&lt;br /&gt;
Position Overview: &lt;br /&gt;
&lt;br /&gt;
We’re looking for a web tier/UI developer&lt;br /&gt;
with demonstrable experience developing high-performance dynamic user interfaces for consumer-facing web sites.&lt;br /&gt;
The ideal candidate will have the skills to develop, enhance and maintain a Rich Internet Application (RIA). &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Position Requirements:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* Expertise using JavaScript and AJAX to develop highly interactive, responsive user interfaces.&lt;br /&gt;
* Extensive knowledge of jQuery and JSON.&lt;br /&gt;
* Working knowledge of Java, JSP and XML.&lt;br /&gt;
* Excellent written and verbal communications skills.&lt;br /&gt;
* Experience in CSS.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Education Requirements:&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
* BS CS/EE or equivalent technical degree with a minimum of 3 years real-world experience.&lt;br /&gt;
&lt;br /&gt;
Think about it.&lt;br /&gt;
One of the world’s largest companies, committed to education and information,&lt;br /&gt;
is on the leading edge of introducing a meaning based search and query tool.&lt;br /&gt;
Indeed we have a platform off of which to work, and a beachhead in the market.&lt;br /&gt;
If you have your sights set high – you should follow up with this one.&lt;br /&gt;
&lt;br /&gt;
Only highly motivated and very smart folks need apply to Susan Blank at susankb48@msn.com&lt;br /&gt;
&lt;br /&gt;
==VP, Search and Semantic Technology. Elsevier Labs - New York==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; Mar 19, 2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
VP, Search and Semantic Technology&amp;lt;br&amp;gt;&lt;br /&gt;
Elsevier Labs &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Description&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Apply online to this position.&lt;br /&gt;
https://reedelsevier.taleo.net/careersection/51/jobdetail.ftl?lang=en&amp;amp;job=40524&lt;br /&gt;
&lt;br /&gt;
Elsevier is currently seeking a VP, Search and Semantic Technology&lt;br /&gt;
to identify with the needs of electronic publishing&lt;br /&gt;
and define a search and discovery technology strategy for next generation electronic products.&lt;br /&gt;
This position works closely with Elsevier product management, Labs, Enterprise Architecture,&lt;br /&gt;
and Elsevier divisional strategy&lt;br /&gt;
to ensure that the semantic technologies implement and inform the product vision.&lt;br /&gt;
&lt;br /&gt;
Responsibilities include:&lt;br /&gt;
&lt;br /&gt;
* Test, implement, program and evaluate search engines and partnerships&lt;br /&gt;
* Keep current with cutting edge research in semantic technologies and translate these into practical timelines for Elsevier&lt;br /&gt;
* Analyze implications of integrating new semantic technologies from both a technical perspective and a user/customer perspective.&lt;br /&gt;
* Work with product strategy to inform and help shape step changes in customer value&lt;br /&gt;
* Work with engineering development managers on assigned projects to plan integration of new technologies&lt;br /&gt;
* Define solutions and evaluate trade-offs for epublishing needs as related to search and other semantic capabilities&lt;br /&gt;
* Work with product architects to ensure that the search technologies meets the specific business needs of future products&lt;br /&gt;
* Manage, develop and own high level technical proposals and effort estimates as the initial piece of the overall product process&lt;br /&gt;
* Determine and recommend technical skills needed to complete a project&lt;br /&gt;
* Responsible for the knowledge transfer of information throughout Elsevier and with the engineering product groups regarding technical issues as related to search and semantic technologies&lt;br /&gt;
* Significantly contribute to knowledge base throughout Elsevier in the form of technical talks, white papers and seminars on technology that support the electronic publishing efforts throughout the business units&lt;br /&gt;
* Contribute to overall platform and technology directions for Elsevier with an emphasis on search&lt;br /&gt;
* Provide Consulting role to projects and product groups for the key search-related technology issues for their business&lt;br /&gt;
* Significantly contribute to RE Ventures&#039; evaluations of emerging search and discovery technologies&lt;br /&gt;
* Assist Enterprise Architects in coordination with RE Applied Technology for the planning and introduction of new search-related technologies&lt;br /&gt;
* Key communicator with IT and product management regarding current and future search capabilities&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications&#039;&#039;&#039;&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;GENERAL KNOWLEDGE:&#039;&#039;&#039;&lt;br /&gt;
 &lt;br /&gt;
* Ability to influence&lt;br /&gt;
* Proven ability to communicate (written and verbal) technology effectively to senior management, product marketing and to engineers&lt;br /&gt;
* Proven ability to absorb large amounts of technical and business detail and synthesize that into a usable problem definition and technical approach&lt;br /&gt;
* Ability to work collaboratively, by directing and guiding the technical direction of a project&lt;br /&gt;
* Ability to work with Senior level development managers&lt;br /&gt;
* Ability to work with ambiguous situations and bring them to closure&lt;br /&gt;
* Ability to influence without direct management&lt;br /&gt;
* Exceptional written and oral communication and presentation skills&lt;br /&gt;
* Demonstrated leadership skills&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;TECHNICAL SKILLS:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Practical and theoretical experience about information retrieval systems&lt;br /&gt;
* Specific technical expertise in these areas:&lt;br /&gt;
* Concept classification, XML, vector space searching, probabilistic retrieval and neural networks&lt;br /&gt;
* Programming techniques: parsing, syntactic analysis, semantic analysis, use of thesauri and ontologies&lt;br /&gt;
* Query processing and query understanding including question/answer paradigms, question extraction, multiple-constraint search paradigms&lt;br /&gt;
* Ability to work with large and diverse data sets&lt;br /&gt;
* Interoperability of search across multiple engines and sites (federated or meta-search) including query normalization&lt;br /&gt;
* PhD in Information Retrieval highly desirable&lt;br /&gt;
* Masters degree in computer science and/or 7+ years of relevant experience in engineering; 5+ years in internet/web technologies&lt;br /&gt;
* Experience as architect (or design lead) in a significant project&lt;br /&gt;
* Experience in transitioning research and prototypes into production a plus&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Other Locations:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* United States-Pennsylvania-Philadelphia&lt;br /&gt;
* United States-Maryland-Rockville&lt;br /&gt;
* United States-California-Irvine&lt;br /&gt;
* United States-Missouri-St Louis&lt;br /&gt;
* United States-New Jersey-Bridgewater&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Closing Date:&#039;&#039;&#039; Ongoing&lt;br /&gt;
&lt;br /&gt;
https://reedelsevier.taleo.net/careersection/51/jobdetail.ftl?lang=en&amp;amp;job=40524&lt;br /&gt;
&lt;br /&gt;
== Funded Startup is Looking for a Experienced Java Engineer ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 02-25-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
AdaptiveBlue (http://getglue.com) is an innovative, well funded semantic web startup in New York City,&lt;br /&gt;
named among the 250 best startups around the world by AlwaysOn Network.&lt;br /&gt;
AdaptiveBlue is focused on results in a fast-paced environment.&lt;br /&gt;
&lt;br /&gt;
We are looking for talented, smart, hard working, experienced and passionate Java Engineer&lt;br /&gt;
to help us build the next generation of web browsing technologies.&lt;br /&gt;
&lt;br /&gt;
You will be working on AdaptiveBlue&#039;s back end - web service, database, and metrics.&lt;br /&gt;
This is an exciting opportunity for someone interested in the semantic technologies,&lt;br /&gt;
skilled with algorithms, and proficient in Java.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Must love coding and hard work&lt;br /&gt;
* B.S. or higher in CS, Math or engineering&lt;br /&gt;
* At lest 5 years of experience as a software engineer&lt;br /&gt;
* At least 5 years of experience in Java programming&lt;br /&gt;
* At least 3 years of experience with SQL&lt;br /&gt;
* Strong knowledge of basic data structures and algorithms&lt;br /&gt;
* Strong knowledge of design patterns, refactoring and unit testing&lt;br /&gt;
* Strong knowledge of concurrent programming&lt;br /&gt;
* Understanding of distributed, large-scale systems&lt;br /&gt;
* Experience with XML/XSL and REST-based Web Services&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;The following are a plus, but not required:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Experience with Amazon Web Services&lt;br /&gt;
* Knowledge of semantic markups and ontologies&lt;br /&gt;
* Experience with semantic APIs like Calais&lt;br /&gt;
* Experience with SQL query tuning&lt;br /&gt;
&lt;br /&gt;
We offer competitive salary, full benefits, 401k, Awesome Mac hardware,&lt;br /&gt;
Herman Miller chairs, and Fresh Direct snacks.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;If you love coding and want to work on exciting things&lt;br /&gt;
that change the way people interact with the web, please send us all of the items below:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* A cover letter describing why you are a good fit&lt;br /&gt;
* A resume with your experiences&lt;br /&gt;
* A sample of Java code that you have written in the past year&lt;br /&gt;
&lt;br /&gt;
* Compensation: Solid base + 401k + stock options&lt;br /&gt;
* Principals only. Recruiters, please don&#039;t contact this job poster.&lt;br /&gt;
* Please, no phone calls about this job!&lt;br /&gt;
* Please do not contact job poster about other services, products or commercial interests.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; job-1049787041@craigslist.org&lt;br /&gt;
&lt;br /&gt;
== Full-time Permanent Semantic Java Software Engineer ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 02-24-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Category:&#039;&#039;&#039; &lt;br /&gt;
Semantic Java Software Engineer&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; &lt;br /&gt;
Boston, MA,  USA - Metro/West&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039; &lt;br /&gt;
We are seeking a full-time permanent semantic java software engineer for our client in the Metro-Boston area.&lt;br /&gt;
If you are interested in building a semantic engine -&lt;br /&gt;
do you consider yourself an ontology expert?&lt;br /&gt;
Know RDF and/or OWL inside and out?&lt;br /&gt;
We&#039;ve got a cool project converting a .Net platform into a Java platform.&lt;br /&gt;
We need a strong coder...&lt;br /&gt;
someone who really really enjoys coding with commercial products experience and a great attitude.&lt;br /&gt;
This you? Give us a call.&lt;br /&gt;
 &lt;br /&gt;
&#039;&#039;&#039;Salary:&#039;&#039;&#039; Open depending on experience&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; &lt;br /&gt;
Sarah Bevelaqua|Technical Recruiter&amp;lt;br&amp;gt;&lt;br /&gt;
The FootBridge Companies |Direct Placement Services Group&amp;lt;br&amp;gt;&lt;br /&gt;
http://www.FootBridgeDirect.com&amp;lt;br&amp;gt;&lt;br /&gt;
Office: 978.474.4455| Toll Free: 877.807.8400&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Artificial Intelligence Software Developer, Washington DC Area==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 02-10-2009&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Category:&#039;&#039;&#039; Artificial Intelligence Software Developer&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Washington, DC, USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Clearance:&#039;&#039;&#039; US DoD Secret required.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Experience:&#039;&#039;&#039; Five years experience in Artificial Intelligence related programming&lt;br /&gt;
or a Master&#039;s Degree in Artificial Intelligence or a related field. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills:&#039;&#039;&#039; Experience in the following, or similar, subjects a plus:&lt;br /&gt;
Software Patterns, Agent-Based Simulation, Software Optimization and Scalability,&lt;br /&gt;
Open Source Contributions, Ontologies, Inference Engines, Evolutionary Computation,&lt;br /&gt;
Neural Networks, Bayesian Networks, Fuzzy Expert Systems, Data Mining,&lt;br /&gt;
Case-Based Reasoning, Game Trees and Game Theory, Social Network Analysis,&lt;br /&gt;
Statistical Design of Experiments. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Programming Language:&#039;&#039;&#039; Java&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Other Software and Development Tools:&#039;&#039;&#039;&lt;br /&gt;
Experience in the following, or similar, software a plus -&lt;br /&gt;
Protégé, Owl, Pellet, Jena, Jastor, Weka, Repast, Groovy, ECJ, Ptolemy,&lt;br /&gt;
JFuzzyLogic, Joone, Jung, R, Jboss, Prefuse&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;&lt;br /&gt;
Artificial Intelligence Software Developers needed&lt;br /&gt;
for cutting-edge Computational Social Science simulation project.&lt;br /&gt;
Will be working in a Team-based environment with other SW developers.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Potential Task areas include:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Scientific Experimentation and Data Mining:  Enhance existing software to support finding patterns in wargame events, for the purposes of testing hypotheses about relations between events and for discovering relationships between events.   Incorporate open source data mining and artificial intelligence software, such as, for example, Weka and ECJ, to automatically find patterns in moves and outcomes&lt;br /&gt;
&lt;br /&gt;
* Interoperation of Hybrid Models:  Support the interoperation of hybrid models, to ensure that models that have multiple resolutions and perspectives share meaning.   Through xml, implement the translations between models.  Integrate with integration tools that support semantic interoperation through ontologies, and other methodologies such as, for example, the COMPOEX backplane or Ptolemy. Help implement a Hub and Spoke design for translation between data models a system by which simulation models and data of different data models may interoperate through their own individual data models, a common data model, and a translation data model between their own individual data models and the common data model. Enhance the software enforcement of a data model &amp;quot;contract&amp;quot; using ontology-based software engineering techniques.  Extend an existing example implemented in Jena and Jastor, converting it to a Dynamic Object Model (DOM) language (such as, for example, Groovy).&lt;br /&gt;
&lt;br /&gt;
* Support for Conflict Resolution between models.  Support the conflict resolution of possibly conflicting hybrid models, for the purposes of making a unified, coherent picture of the social environment.&lt;br /&gt;
Integrate open source software (such as, for example, JFuzzyLogic) to match the simulation output to correlative social study data, and establish the correlative relations that should exist between and within component social science models in support of validation and consensus building of possibly conflicting models.&lt;br /&gt;
Implement a framework for consensus-building.  Make possible the specification of arbitrary schemes for developing a model consensus.  The framework would have a way to handle issues of integration, for example, the models may conflict by having mutually exclusive or uncorrelated behaviors.   The framework would allow various consensus schemes to be switched in and out.  Implement a simple prototype weighted voting scheme in the framework&lt;br /&gt;
&lt;br /&gt;
* Support for Component Models.  Enhance the Nexus Intelligent Adaptive Agent Based Models to increase their generality, scalability, efficiency, and ability to work as component models for other software.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; &amp;lt;br&amp;gt;&lt;br /&gt;
Debbie Duong &amp;lt;br&amp;gt;&lt;br /&gt;
debbieduong62  at gmail.com&lt;br /&gt;
&lt;br /&gt;
==Knight Professor of the Practice of Journalism and Public Policy Studies, Duke University==&lt;br /&gt;
&lt;br /&gt;
Duke University&amp;lt;br&amp;gt;&lt;br /&gt;
Sanford Institute of Public Policy&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Durham, NC, USA&lt;br /&gt;
&lt;br /&gt;
The Sanford Institute of Public Policy at Duke University seeks applicants&lt;br /&gt;
for the Knight Professor of the Practice of Journalism and Public Policy Studies,&lt;br /&gt;
an endowed chair in Duke’s DeWitt Wallace Center for Media and Democracy.&lt;br /&gt;
&lt;br /&gt;
We seek a person who will help in the development of a new field, computational journalism.&lt;br /&gt;
Advances in data availability, technology, and algorithms offer the prospect&lt;br /&gt;
that part of the media’s watchdog function may be supplemented by analyses done by computers.&lt;br /&gt;
The Knight Professor at Duke will help translate advances&lt;br /&gt;
in areas such as artificial intelligence and the semantic web&lt;br /&gt;
into the development of products that help journalists and other community members&lt;br /&gt;
hold institutions accountable.&lt;br /&gt;
In part, this may involve computerizing aspects of investigative reporting.&lt;br /&gt;
&lt;br /&gt;
Candidates for the chair should have experience in working&lt;br /&gt;
at the intersection of technology and information generation.&lt;br /&gt;
They may be working in computer assisted reporting, or visualization of data on the Web,&lt;br /&gt;
or the development of algorithms to sift through publicly available data&lt;br /&gt;
for clues to the performance of government.&lt;br /&gt;
Ideal candidates could include reporters or editors involved in innovations in digital news ventures&lt;br /&gt;
or programmers involved in the development of artificial intelligence or semantic web products&lt;br /&gt;
aimed at news gathering and reporting.&lt;br /&gt;
Candidates should be strongly committed to advancing the development of computational journalism as a field&lt;br /&gt;
through innovative research and the development of new digital tools.&lt;br /&gt;
&lt;br /&gt;
The Knight Professor will teach courses in computational journalism and public policy&lt;br /&gt;
in our undergraduate public policy major and certificate program in policy journalism and media studies.&lt;br /&gt;
The Knight Professor will play a key role in the teaching, research and policy engagement activities&lt;br /&gt;
of the DeWitt Wallace Center for Media and Democracy.&lt;br /&gt;
&lt;br /&gt;
Candidates for this position should send a CV and other materials to:&lt;br /&gt;
&lt;br /&gt;
Professor James T. Hamilton&amp;lt;br&amp;gt;&lt;br /&gt;
Knight Search Committee Chair&amp;lt;br&amp;gt;&lt;br /&gt;
DeWitt Wallace Center for Media and Democracy&amp;lt;br&amp;gt;&lt;br /&gt;
Sanford Institute of Public Policy&amp;lt;br&amp;gt;&lt;br /&gt;
Duke University&amp;lt;br&amp;gt;&lt;br /&gt;
Box 90241&amp;lt;br&amp;gt;&lt;br /&gt;
Durham, NC 27708-0241&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Review of applications will begin immediately.&lt;br /&gt;
Applications received by December 31, 2008 will be guaranteed full consideration.&lt;br /&gt;
&lt;br /&gt;
Duke University is an Equal Opportunity/Affirmative Action Employer.&lt;br /&gt;
&lt;br /&gt;
==Executive Director Emerging Platforms ORGANIC, INC.==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 10-23-08&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Area code:&#039;&#039;&#039; 212&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Pay rate:&#039;&#039;&#039; open&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Experience:&#039;&#039;&#039; Exceptional Experience.&lt;br /&gt;
&lt;br /&gt;
Organic is all about the exceptional.&lt;br /&gt;
We are a leading digital communications agency –&lt;br /&gt;
the first, in fact – focused on designing and building exceptional experiences&lt;br /&gt;
that help make the online channel really work for leading companies.&lt;br /&gt;
We work on everything digital – from websites to online marketing campaigns,&lt;br /&gt;
from mobile applications to digital billboards –&lt;br /&gt;
for companies such as Chrysler LLC, Warner Bros. International, Geek Squad and Bank of America.&lt;br /&gt;
We are passionate about emerging platforms,&lt;br /&gt;
and this helps keep us on the edge of the hottest and most interesting stories in the industry.&lt;br /&gt;
And, as a leader in the marketplace, we have the opportunity to work on really exciting projects. &lt;br /&gt;
&lt;br /&gt;
Organic has an open, diverse culture that encourages participation and innovation.&lt;br /&gt;
We reward exceptional work and creative ideas.&lt;br /&gt;
At Organic we believe in working hard and having fun while we’re doing it.&lt;br /&gt;
We have great benefits and perks including massages, pet insurance, Wednesday bagels and Friday cocktails. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* The Executive Director of Emerging Platforms leads client innovation in the areas of new media technologies, platforms and networks. &lt;br /&gt;
* Oversees the Emerging Platform team which is comprised of strategist and interactive technology developers. &lt;br /&gt;
* Provides clients with an outward looking perspective of the new media landscape and demonstrates what clients should expect and how these developments may impact their online experience, business, industry, competition and customers. &lt;br /&gt;
* Leads the rapid prototyping initiative for Organic. &lt;br /&gt;
* Develop internal projects and employee education programs to demonstrate new technologies to Organic employees with the hope of eventual client education and suggestion. &lt;br /&gt;
* Builds and manages connections to the new media industry to identify new opportunities for Organic clients and Organic employees (toolsets, frameworks and development technologies). &lt;br /&gt;
* Industry perspective and thought leadership platforms for Organic and for industry media (corporate marketing). &lt;br /&gt;
* Attendance and Organic&#039;s representation at industry conferences and events. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Education and Work Experience:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* 10 yrs overall experience preferred including: &lt;br /&gt;
* 10+ yrs experience working in marketing strategy, account planning and interactive marketing. &lt;br /&gt;
* Experience in building, growing and leading successful teams (either in an office, practice, or global teams of 50+). &lt;br /&gt;
* Client-focused – builds long term high-level relationships that reflect well on Organic, proven history in business management and growth. &lt;br /&gt;
* Ability to manage and be held accountable for resources, revenues and budgets. &lt;br /&gt;
* Understands and can both, manage and work within a matrix management organization. &lt;br /&gt;
* Personable; very poised; a strong conceptual thinker; have high energy; exhibit excellent writing, communication and presentation skills. &lt;br /&gt;
* Inspirational leader – exhibits integrity, and respectful interactions with past recognizable clients, superiors, peers and subordinates. &lt;br /&gt;
* Proven history of growing a practice or major service. &lt;br /&gt;
* Able to thrive in a &amp;quot;non-traditional&amp;quot;, entrepreneurial environment. &lt;br /&gt;
* Availability and willingness to travel 35%of time. &lt;br /&gt;
* Bachelors/MBA degree preferred or equivalent work experience. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Travel required:&#039;&#039;&#039; 35%&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Telecommute:&#039;&#039;&#039; no&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039; Ross, http://www.organic.com, ross @ organic . com&lt;br /&gt;
&lt;br /&gt;
==Web 2.0 Build Out==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 10-22-2008&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
What if there was absolutely nothing stopping you from inventing the next big thing?&lt;br /&gt;
What would you do if you had a clear voice in developing the next generation of web based information services?&lt;br /&gt;
How would you respond if you had the opportunity to be a creative force&lt;br /&gt;
in one of the most unique and inspiring work cultures in NYC? &lt;br /&gt;
&lt;br /&gt;
Our client is a startup web unit within a global telecom leader,&lt;br /&gt;
launching a mobile and web based platform&lt;br /&gt;
that will revolutionize the way consumers interact with and manage information.&lt;br /&gt;
They are early in development and building a team of innovative, intelligent professionals who love what they do.&lt;br /&gt;
We are looking for people who are obsessed with web technology, creative in their process&lt;br /&gt;
and can thrive in an entrepreneurial, fast paced environment.&lt;br /&gt;
Our ideal candidates work very well within a team, are self-motivated,&lt;br /&gt;
have a high level of energy and a strong drive to succeed.&lt;br /&gt;
They must have a passion for the internet and be current with cutting edge web trends and technologies.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Primary Responsibilities:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Work as key member of a small team developing foundational systems and services for an outstanding web 2.0 platform &lt;br /&gt;
* Work with vendors and agencies to integrate products and services into platform &lt;br /&gt;
* Hands on development of web systems and projects &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Development experience in standard internet technologies &lt;br /&gt;
* Experience in an SOA environment is preferred &lt;br /&gt;
* Linux deployment &lt;br /&gt;
* Solid Python development skills &lt;br /&gt;
* System design experience preferred &lt;br /&gt;
* Track record of successful releases &lt;br /&gt;
* Experience with large scale distributed systems &lt;br /&gt;
* RESTful services a plus &lt;br /&gt;
* JavaScripting, JSON, XML, Ruby, Erlang all a plus&lt;br /&gt;
&lt;br /&gt;
Interested candidates email resumes to jim@execuseek.net&lt;br /&gt;
&lt;br /&gt;
==Program Manager with Ontology and Taxonomy/OWL==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Skills:&#039;&#039;&#039; Taxonomy Ontology tools OWL+ Media Endeca&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; 7-1-2008&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Area code:&#039;&#039;&#039; 212&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Pay rate:&#039;&#039;&#039; open&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Job Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
This individual will be responsible for multiple aspects of simultaneous long term projects including:&lt;br /&gt;
Consolidating project plans, responsibility for deadlines, ongoing support,&lt;br /&gt;
implementation and coordination and overall delivery of multiple inter-related projects.&lt;br /&gt;
&lt;br /&gt;
Accountable for the on time and on budget completion of these projects,&lt;br /&gt;
the Program Manager will define, coordinate and lead a blended team&lt;br /&gt;
thus requiring regular interaction with business, editorial, production, marketing, design and development factions&lt;br /&gt;
of the internet organizations.&lt;br /&gt;
The individual is expected to regularly participate in all phases of project development&lt;br /&gt;
with a strong team-oriented attitude.&lt;br /&gt;
Strong communication skills, both written and verbal are essential,&lt;br /&gt;
as is the ability to simultaneously manage multiple projects in a dynamic, challenging, and fast-paced environment.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Candidates will possess a Bachelor&#039;s Degree in a related field&lt;br /&gt;
* 5-7 years Project Management Experience with solid experience working in a cross-functional environment.&lt;br /&gt;
* This person will be a **manager of project managers** and should have experience pulling together very large programs of projects with multiple large interrelated subprojects&lt;br /&gt;
* Experience with search technologies is a must; specific experience with Endeca a major plus&lt;br /&gt;
* Experience with taxonomy / ontology tools a must; specific experience with OWL-based technologies a major plus&lt;br /&gt;
* Requires a high-level understanding of the integration of search and taxonomy/ontology tools with a content management system&lt;br /&gt;
* Online Publishing / Media experience a huge plus&lt;br /&gt;
* GREAT interpersonal skills&lt;br /&gt;
** this person will interface with the project managers, functional managers, and will be required to have frequent written and face-to-face communications with senior management&lt;br /&gt;
** Managing software vendor relationships and coordinates training, documentation, and communication with IT and the Titles.&lt;br /&gt;
* Creates and executes project work plans and revises as appropriate to meet changing needs and requirements.&lt;br /&gt;
* Creates and executes an aggregate program view of interrelated projects and has the ability to recognize, surface, and resolve issues across the program.&lt;br /&gt;
* PMI Certification is preferred, but not mandatory.&lt;br /&gt;
* Experienced with Web based projects, preferably in the online content/publishing space.&lt;br /&gt;
* Effectively applies our methodology and enforces project standards with an eye on emerging industry practices.&lt;br /&gt;
* Prepares for engagement reviews and quality assurance procedures.&lt;br /&gt;
* Recognizes and minimizes exposures and risks on the individual projects and across the program of projects.&lt;br /&gt;
* Ensures project documents are complete, current, and stored appropriately.&lt;br /&gt;
* Facilitates team and client meetings effectively, documents meeting minutes and agreements.&lt;br /&gt;
* Effectively communicates relevant project information to superiors in an engaging, informative, well-organized format. Must be able to effectively tailor communications of complex issues to audiences with varying levels of subject matter expertise.&lt;br /&gt;
* Acquires / possesses a thorough understanding of our capabilities.&lt;br /&gt;
* Resolves and/or escalates issues in a timely fashion and understands how to communicate difficult/sensitive information tactfully.&lt;br /&gt;
* Motivates team to work together in the most efficient manner.&lt;br /&gt;
* Suggests areas for improvement in internal processes along with possible solutions.&lt;br /&gt;
* Expert in MS Project, MS Excel, MS Powerpoint&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Travel required:&#039;&#039;&#039; none&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Telecommute:&#039;&#039;&#039; no&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Debbie Levy, [http://www.logiccorporation.com Logic]&lt;br /&gt;
&amp;lt;br&amp;gt;debbie @ logiccorporation . com&lt;br /&gt;
&lt;br /&gt;
== Semantic Web Senior Java Developer -- Alitora Systems ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; New York, NY, USA&lt;br /&gt;
&lt;br /&gt;
Alitora Systems is a start-up software company providing an innovative Semantic Search and Collaboration service.&lt;br /&gt;
&lt;br /&gt;
See: [http://www.alitora.com http://www.alitora.com]&lt;br /&gt;
&lt;br /&gt;
We are building the first true commercial Semantic Web Service as a SaaS platform.&lt;br /&gt;
We are primarily serving the Life Science Industry, such as the pharmaceutical industry,&lt;br /&gt;
which requires a huge amount of critical high value data&lt;br /&gt;
to support Drug Development, Competitive Intelligence, and Business Development.&lt;br /&gt;
&lt;br /&gt;
Our core technology is the kHarmony Semantic Database:&lt;br /&gt;
a general hyper-graph database used to store knowledge modeled semantically.&lt;br /&gt;
&lt;br /&gt;
We are seeking a senior developer with significant Java experience, including:&lt;br /&gt;
&lt;br /&gt;
* Strong Experience with Semantic technologies and/or Search Engine technologies, such as:&lt;br /&gt;
**Java Lucene search engine&lt;br /&gt;
**Machine Learning algorithms and toolkits&lt;br /&gt;
**JENA and related Java toolkits&lt;br /&gt;
&lt;br /&gt;
* Significant Experience implementing efficient algorithms in areas such as:&lt;br /&gt;
** Search engine query processing and/or index construction&lt;br /&gt;
** Inference algorithms, Logic Algorithms&lt;br /&gt;
** Machine Learning / Statistical Methods&lt;br /&gt;
** Graph Theoretic Algorithms&lt;br /&gt;
** Natural Language Processing algorithms, such as Entity Extraction&lt;br /&gt;
&lt;br /&gt;
* Understanding of Semantic Web concepts, such as:&lt;br /&gt;
** RDF, OWL, Inference Engines, URIs&lt;br /&gt;
&lt;br /&gt;
* Understanding of Graph Theory concepts, such as:&lt;br /&gt;
** Nodes, Edges, Depth-First-Search, Cliques&lt;br /&gt;
&lt;br /&gt;
* Appreciation of the Life Sciences&lt;br /&gt;
&lt;br /&gt;
* Great inter-personal / teamwork skills and communication skills&lt;br /&gt;
&lt;br /&gt;
* Able to prosper in a cutting-edge start-up environment&lt;br /&gt;
&lt;br /&gt;
* And especially, Passionate Enthusiasm for New Technology and Start-Ups&lt;br /&gt;
&lt;br /&gt;
---------------&lt;br /&gt;
&lt;br /&gt;
Please reply with a resume.&lt;br /&gt;
&lt;br /&gt;
We’re seeking to bring developers on as consultants or employees, depending on individual circumstance&lt;br /&gt;
&lt;br /&gt;
We&#039;re based in New York City, although some developers work remotely. We have daily meetings EST.&lt;br /&gt;
&lt;br /&gt;
Email: &#039;&#039;marc@alitora.com&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
==Experienced Java Programmer for Semantic R&amp;amp;D Position (Washington, DC)==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date: ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Washington, DC, USA&lt;br /&gt;
&lt;br /&gt;
Small, established startup (http://clarkparsia.com/) in downtown DC&lt;br /&gt;
seeks an experienced, systems-level Java programmer&lt;br /&gt;
to work on R&amp;amp;D projects and production systems in semantic technologies,&lt;br /&gt;
including reasoning, planning, description logics, semantic web services, logistics, etc.&lt;br /&gt;
Join the team responsible for Pellet, the leading OWL DL reasoner.&lt;br /&gt;
&lt;br /&gt;
The successful applicant will be smart and have a demonstrable record of getting things done.&lt;br /&gt;
&lt;br /&gt;
* Bachelor&#039;s or Master&#039;s in CS or relevant field&lt;br /&gt;
* Minimum 7 years of serious programming experience&lt;br /&gt;
* Experience with AI, KR, logic programming, or planning a definite plus&lt;br /&gt;
* Enthusiasm and ability for solving difficult, algorithmically novel problems required&lt;br /&gt;
* Familiarity with database theory valuable&lt;br /&gt;
* Familiarity with systems or operations research also valuable&lt;br /&gt;
* Excellent written and verbal communication skills&lt;br /&gt;
* Must be eligible to work permanently in the US&lt;br /&gt;
&lt;br /&gt;
We offer the following benefits:&lt;br /&gt;
&lt;br /&gt;
* Competitive salary and benefits (401k, FSA, major med, dental, vision, etc)&lt;br /&gt;
* A chance to build new, cool stuff that people use&lt;br /&gt;
* Interesting, vibrant work environment: &lt;br /&gt;
** Free lunch, &lt;br /&gt;
** baseball and other sports outings, &lt;br /&gt;
** Wii tournaments,&lt;br /&gt;
** great espresso &amp;amp; coffee &lt;br /&gt;
** Beer Fridays&lt;br /&gt;
** 1 block from Metro (Mt Vernon Square, Green Line) near Convention Center&lt;br /&gt;
&lt;br /&gt;
To apply, send resume and a cover letter to Kendall Clark, kendall@clarkparsia.com.&lt;br /&gt;
&lt;br /&gt;
==Ontology Modeler Systems Engineer, Mountain View CA==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Mountain View, CA, USA&lt;br /&gt;
&lt;br /&gt;
My client is in Mountain View CA.&lt;br /&gt;
They are looking for an Ontology Modeler/Systems Engineer.&lt;br /&gt;
The primary focus will be on developing information models for Aerospace Vehicles and Systems.&lt;br /&gt;
This includes general enterprise architecture modeling (e.g., organizations, processes, tools)&lt;br /&gt;
as well as models specific to engineering domains (e.g., vehicles, sub-systems, devices and functions).&lt;br /&gt;
You will be using Semantic Web standards (RDF and OWL) and XML.&lt;br /&gt;
Any knowledge of these technologies is a big plus, but we will train the right person.  &lt;br /&gt;
&lt;br /&gt;
This position requires:&lt;br /&gt;
&lt;br /&gt;
* More than 5 years of modeling experience (either object modeling, system modeling, data modeling or knowledge modeling)&lt;br /&gt;
* A degree and/or work experience in an engineering field&lt;br /&gt;
* Strong communications skills (experience interviewing people for knowledge capture, running workshops, etc.)&lt;br /&gt;
&lt;br /&gt;
Knowledge of space systems and their engineering disciplines is a distinct advantage.&lt;br /&gt;
This includes avionics, mechanics, hydraulics, propulsion, guidance and navigation,&lt;br /&gt;
telemetry and control systems. Some past or current programming skills will be an advantage,&lt;br /&gt;
especially in Java, Prolog, or a functional language such as Haskell.&lt;br /&gt;
&lt;br /&gt;
The ideal candidate would also have one (or more) of the following qualifications.&lt;br /&gt;
Knowledge and experience of:&lt;br /&gt;
&lt;br /&gt;
* Modeling formalisms such as UML and SysML.&lt;br /&gt;
* Information and knowledge structuring formalisms such as ASN.1, XML, RDF, OWL.&lt;br /&gt;
* Enterprise Architecture frameworks such as DODAF and TOGAF&lt;br /&gt;
* Training and/or experience in computational  linguistics&lt;br /&gt;
&lt;br /&gt;
The person should enjoy working on challenging problems, be a self starter,&lt;br /&gt;
have strong communication skills and be ready to show a lot of initiative.&lt;br /&gt;
He/she should want to work in a dynamic and growing small company&lt;br /&gt;
with collegial culture and many opportunities to learn and do different things. &lt;br /&gt;
&lt;br /&gt;
We offer competitive salary, bonus, major medical, dental, and an attractive stock option plan.&lt;br /&gt;
The position is in our Mountain View, California location.&lt;br /&gt;
We will provide assistance with the relocation expenses.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Mary Frances Hunter&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
503 232 8822&amp;lt;br&amp;gt;&lt;br /&gt;
&#039;&#039;Recruiter Extraordinaire&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Software Engineer, Mountain View CA==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Mountain View, CA, USA&lt;br /&gt;
&lt;br /&gt;
My client is in Mountain View CA. They are looking for Software Engineers who know server-side Java development and want to work on a Semantic Web product.&lt;br /&gt;
&lt;br /&gt;
The Product Suite is built on Eclipse but with a Web UI. They want to add more features. This involves design, development, test and integration of the new features. We need someone who is a self-starter, as there is little supervision. There is collaboration.&lt;br /&gt;
&lt;br /&gt;
New feature may include the importing and exporting of data, merging and transferring data, connecting to external data bases, managing changes to forms…..&lt;br /&gt;
&lt;br /&gt;
Technologies involved are: Java, Eclipse, Topcat, Adobe Products, AJAX/Flex, Mash-ups, OWL, RDF/S, SPARQL… &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Mary Frances Hunter&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
503 232 8822&amp;lt;br&amp;gt;&lt;br /&gt;
&#039;&#039;Recruiter Extraordinaire&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Senior Engineer, Semantic Web Datastore ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; ?&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Cambridge, MA, USA&lt;br /&gt;
&lt;br /&gt;
ITA Software&lt;br /&gt;
&lt;br /&gt;
http://www.itasoftware.com/careers/jlisting.html?jid=26&lt;br /&gt;
&lt;br /&gt;
ITA Software -- known for its algorithm-intensive airfare search and reservations products --&lt;br /&gt;
has also been doing novel Semantic Web and Data Integration work since 2004.&lt;br /&gt;
We&#039;re now turning the corner from research to deployment&lt;br /&gt;
and seek the right person to engineer our data store for scalability and performance.&lt;br /&gt;
Follow the link above for details.&lt;br /&gt;
&lt;br /&gt;
[Thanks, Marco, for inviting me to post here!  -JustinITA]&lt;br /&gt;
&lt;br /&gt;
=Archive=&lt;br /&gt;
&lt;br /&gt;
==Information Architect, Collection Information &amp;amp; Access, J. Paul Getty Museum==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Date:&#039;&#039;&#039; December 2008&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Location:&#039;&#039;&#039; Los Angeles, CA, USA&lt;br /&gt;
&#039;&#039;&#039;Description:&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
The department of Collection Information &amp;amp; Access at the J. Paul Getty Museum&lt;br /&gt;
is seeking an Information Architect to oversee the back-end structure, data models, systems and applications&lt;br /&gt;
used by the Museum to support the management and dissemination of documentation, digital assets, and metadata&lt;br /&gt;
on the collection and to ensure its accessibility in the networked environment.&lt;br /&gt;
The Information Architect will lead efforts in restructuring the way information&lt;br /&gt;
is stored, systems integrated, and data published&lt;br /&gt;
so as best to ensure efficiency in processes, scalability and sustainability, and resource discovery.&lt;br /&gt;
This will involve architectural designs, analysis, integration, and strategic direction&lt;br /&gt;
for how best to manage existing enterprise-wide applications&lt;br /&gt;
such as collections management, content management and digital asset management systems&lt;br /&gt;
with other custom grown applications and open source solutions,&lt;br /&gt;
in addition to overseeing data modeling and strategies that are system independent.&lt;br /&gt;
The position will be responsible for the maintenance of data models, data dictionaries, and processes;&lt;br /&gt;
work with technical staff across the Getty to build mechanisms&lt;br /&gt;
for exchanging data and metadata between repositories;&lt;br /&gt;
and work closely with user communities&lt;br /&gt;
for requirements analysis, problem definition and solutions development. &lt;br /&gt;
&lt;br /&gt;
The ideal candidate will utilize standards, best practices, and forward-thinking solutions&lt;br /&gt;
for structuring the Museum&#039;s information architecture,&lt;br /&gt;
and be able to provide analysis, documentation, and ROI for strategies.&lt;br /&gt;
The candidate should have experience in all phases of the software development cycle;&lt;br /&gt;
understand and be technically proficient in the environments&lt;br /&gt;
in which software applications operate (i.e. Unix, Windows);&lt;br /&gt;
have familiarity with semantic technologies&lt;br /&gt;
including triple stores, natural language processing, and clustering techniques.&lt;br /&gt;
The candidate should be comfortable with writing technical documentation and design documents,&lt;br /&gt;
outlining detailed process flow and workflow mappings, have strong analytical skills,&lt;br /&gt;
excellent oral and written communication skills,&lt;br /&gt;
and the ability to effectively work in a team environment. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Requirements:&#039;&#039;&#039; Proven experience working with relational databases (Oracle 10g), SQL Server and using Structured Query Language; familiarity with &amp;quot;C++&amp;quot;, JAVA or similar object-oriented programming language; and proficient at UNIX scripting languages; JavaScript, HTML, CSS, XML and XSLT. Working knowledge of ontologies and ontology standards like RDF and concepts associated with the Semantic Web.  &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Qualifications:&#039;&#039;&#039; Bachelor&#039;s Degree in Computer Science, Library &amp;amp; Information Science, Information Technology, or related studies required, Master&#039;s preferred. Minimum 8 years of experience in the electronic management of information, and developing, implementing and managing information architecture in a publishing, library, or educational repository environment strongly preferred. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Contact:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
Please email cover letter and resume to jobs@getty.edu indicating in the subject line, &amp;quot;Museum Information Architect /AT.&amp;quot; OR send to: The J. Paul Getty Trust, 1200 Getty Center Drive, Suite 400, Los Angeles, Ca 90049-1681 and reference &amp;quot; Museum Information Architect /AT&amp;quot; in your cover letter. No phone calls, please. EOE.&lt;br /&gt;
&lt;br /&gt;
http://www.getty.edu/about/opportunities/tech_opps.html&lt;br /&gt;
&lt;br /&gt;
== Global Director of Semantic Technology Solutions ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Providing thought-leadership in the adoption of emerging semantic technologies as a means to add value to information products and thereby drive revenue.&lt;br /&gt;
* Overseeing software development of Synaptica® from Dow Jones, an enterprise-class taxonomy and ontology management software product.&lt;br /&gt;
* Leading the development of new taxonomies and metadata that can be used to enrich information content.&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Full Job Description online at:&lt;br /&gt;
&lt;br /&gt;
http://careers.peopleclick.com/jobposts/Client40_DowJones/BU1/External/pck314-6886.htm&lt;br /&gt;
&lt;br /&gt;
Dave Clarke&amp;lt;br&amp;gt;&lt;br /&gt;
Global Taxonomy Director&amp;lt;br&amp;gt;&lt;br /&gt;
Dow Jones&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Global Taxonomy Director ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Coordinate initiatives with the regional Client Solutions Directors to achieve and exceed Dow Jones Taxonomy Services revenue and operational goals, leveraging Dow Jones&#039; taxonomy and editorial expertise.&lt;br /&gt;
* Define and communicate Taxonomy capabilities, services and pricing models.&lt;br /&gt;
* Ensure structures and procedures are in place to achieve the successful delivery of taxonomy engagements to time, to budget and to satisfactory quality standards.&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Full Job Description online at:&lt;br /&gt;
&lt;br /&gt;
http://careers.peopleclick.com/jobposts/Client40_DowJones/BU1/External/pck314-6888.htm&lt;br /&gt;
&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
Dave Clarke&amp;lt;br&amp;gt;&lt;br /&gt;
Global Taxonomy Director&amp;lt;br&amp;gt;&lt;br /&gt;
Dow Jones&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Start-up Seeking Software Engineer Computational Linguistics  ==&lt;br /&gt;
&lt;br /&gt;
We are a data driven start-up looking for a CTO to build a social media analytics application&lt;br /&gt;
using Natural Language Processing with experience in web-scraping and text mining.&lt;br /&gt;
Candidate will develop methodologies to find sentiment&lt;br /&gt;
(sentiment in context, polarity, intensity, relationships between issues and reasons for issues),&lt;br /&gt;
tuning for less formal content (tokenization, sentence segmentation, part-of-speech (POS) tagging, parsing, etc.)&lt;br /&gt;
and finding embedded meaning and intelligence in less grammatical text.&lt;br /&gt;
Experience in predictive modeling in data mining a plus.&lt;br /&gt;
&lt;br /&gt;
Qualifications:&lt;br /&gt;
* 2-4 years experience in a relevant field such as software engineering, machine learning programming, et al.&lt;br /&gt;
* The drive to work in a fast-paced, multi-disciplinary start-up environment.&lt;br /&gt;
* A strong interest in online social networks and network/user dynamics.&lt;br /&gt;
* PhD in a computational linguistics (or related field) required.&lt;br /&gt;
* Experience in database design and search algorithms.&lt;br /&gt;
* Experience with Python/MySQL/R and SaaS is preferred but not required.&lt;br /&gt;
* Excellent documentation &amp;amp; whitepaper authoring skills.&lt;br /&gt;
* Patent filing experience preferred.&lt;br /&gt;
&lt;br /&gt;
I am looking for an exceptional management team member who will provide design the product vision&lt;br /&gt;
and has extensive experience in delivering complex web-based applications of significant scale&lt;br /&gt;
versed in the social semantic web.&lt;br /&gt;
Feedback of the product framework has been extremely positive;&lt;br /&gt;
we are looking to build a working prototype of site.&lt;br /&gt;
There is no direct compensation for this role.&lt;br /&gt;
Equity stake in company for right candidate.&lt;br /&gt;
This role will become a full time position once funded.&lt;br /&gt;
&lt;br /&gt;
This position will require telecommuting as we are located in Newburyport, MA.&lt;br /&gt;
&lt;br /&gt;
Learn more: http://www.socialtality.com &lt;br /&gt;
&lt;br /&gt;
Contact me directly at wendytroupe [at] socialtality [dot] com&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6651</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6651"/>
		<updated>2026-04-17T15:59:16Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIO&#039;s in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
&lt;br /&gt;
Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
&lt;br /&gt;
It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take.&lt;br /&gt;
&lt;br /&gt;
So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know, he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms. Ontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
&lt;br /&gt;
So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution. So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal. Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
&lt;br /&gt;
Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um, we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
&lt;br /&gt;
So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this? Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes, it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Well you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools. All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies. obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around. Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And  in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information? &amp;quot;They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas. So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh, if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my world is a very broad very shallow world Peter the other well I would say on on the scientific side there&#039;s not an enormous amount of what&#039;s termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise. So these examples I think are more complex understanding and then integration. so I think actually that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t quite map perfectly and then we have to basically do a translation step so that they can actually be integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah. Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So, in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
&lt;br /&gt;
Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide there. And that really is starting to change at the international level and this is true for science and medicine. there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be infrastructure for well I might be giving a control or permission to this company, but now this company has all these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes, but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money? The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies. Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway. Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know, again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there there&#039;s value somewhere in here. And we&#039;re being asked exactly the question, where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be advertising instead of pornography. It surprised a lot of us. Um, you know, some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene. haven&#039;t been there, you should check it out maybe of it. semantic technologies. It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question? Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data. The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because they one person saw in the data and it turned out to be something really of interest. I think actually we can point to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and then when he decided off through his own intuition that this would change the way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
&lt;br /&gt;
I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6650</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6650"/>
		<updated>2026-04-17T15:41:45Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
&lt;br /&gt;
Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
&lt;br /&gt;
It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take.&lt;br /&gt;
&lt;br /&gt;
So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know, he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms. Ontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
&lt;br /&gt;
So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution. So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal. Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
&lt;br /&gt;
Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um, we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
&lt;br /&gt;
So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this? Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes, it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Well you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools. All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies. obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around. Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And  in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information? &amp;quot;They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas. So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh, if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my world is a very broad very shallow world Peter the other well I would say on on the scientific side there&#039;s not an enormous amount of what&#039;s termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise. So these examples I think are more complex understanding and then integration. so I think actually that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t quite map perfectly and then we have to basically do a translation step so that they can actually be integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah. Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So, in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
&lt;br /&gt;
Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide there. And that really is starting to change at the international level and this is true for science and medicine. there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be infrastructure for well I might be giving a control or permission to this company, but now this company has all these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes, but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money? The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies. Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway. Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know, again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there there&#039;s value somewhere in here. And we&#039;re being asked exactly the question, where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be advertising instead of pornography. It surprised a lot of us. Um, you know, some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene. haven&#039;t been there, you should check it out maybe of it. semantic technologies. It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question? Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data. The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because they one person saw in the data and it turned out to be something really of interest. I think actually we can point to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and then when he decided off through his own intuition that this would change the way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
&lt;br /&gt;
I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6649</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6649"/>
		<updated>2026-04-17T15:40:34Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
&lt;br /&gt;
Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
&lt;br /&gt;
It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take.&lt;br /&gt;
&lt;br /&gt;
So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know, he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms. Ontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
&lt;br /&gt;
So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution. So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal. Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
&lt;br /&gt;
Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um, we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
&lt;br /&gt;
So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this? Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes, it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Well you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools. All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies. obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around. Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And  in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information? &amp;quot;They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas. So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh, if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my world is a very broad very shallow world Peter the other well I would say on on the scientific side there&#039;s not an enormous amount of what&#039;s termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise. So these examples I think are more complex understanding and then integration. so I think actually that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t quite map perfectly and then we have to basically do a translation step so that they can actually be integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah. Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So, in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
&lt;br /&gt;
Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide there. And that really is starting to change at the international level and this is true for science and medicine. there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be infrastructure for well I might be giving a control or permission to this company, but now this company has all these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes, but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money? The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies. Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway. Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know, again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there there&#039;s value somewhere in here. And we&#039;re being asked exactly the question, where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be advertising instead of pornography. It surprised a lot of us. Um, you know, some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene. haven&#039;t been there, you should check it out maybe of it. semantic technologies. It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question? Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data. The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because they one person saw in the data and it turned out to be something really of interest. I think actually we can point to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and then when he decided off through his own intuition that this would change the way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
&lt;br /&gt;
I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6648</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6648"/>
		<updated>2026-04-17T15:39:47Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
&lt;br /&gt;
Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
&lt;br /&gt;
It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take.&lt;br /&gt;
&lt;br /&gt;
So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know, he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms. Ontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
&lt;br /&gt;
So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution. So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal. Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
&lt;br /&gt;
Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um, we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
&lt;br /&gt;
So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this? Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes, it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Well you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools. All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies. obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around. Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And  in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information? &amp;quot;They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas. So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh, if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my world is a very broad very shallow world Peter the other well I would say on on the scientific side there&#039;s not an enormous amount of what&#039;s termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise. So these examples I think are more complex understanding and then integration. so I think actually that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t quite map perfectly and then we have to basically do a translation step so that they can actually be integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah. Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So, in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
&lt;br /&gt;
Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide there. And that really is starting to change at the international level and this is true for science and medicine. there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be infrastructure for well I might be giving a control or permission to this company, but now this company has all these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes, but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money? The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies. Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway. Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know, again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there there&#039;s value somewhere in here. And we&#039;re being asked exactly the question, where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be advertising instead of pornography. It surprised a lot of us. Um, you know, some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene. haven&#039;t been there, you should check it out maybe of it. semantic technologies. It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question? Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data. The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because they one person saw in the data and it turned out to be something really of interest. I think actually we can point to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and then when he decided off through his own intuition that this would change the way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
&lt;br /&gt;
I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6647</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6647"/>
		<updated>2026-04-17T15:37:37Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
&lt;br /&gt;
Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
&lt;br /&gt;
It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take.&lt;br /&gt;
&lt;br /&gt;
So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know, he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms. Ontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
&lt;br /&gt;
So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution. So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal. Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
&lt;br /&gt;
Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um, we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
&lt;br /&gt;
So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this? Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes, it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Well you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies. obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around. Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And  in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So, in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
&lt;br /&gt;
Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide there. And that really is starting to change at the international level and this is true for science and medicine. there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be infrastructure for well I might be giving a control or permission to this company, but now this company has all these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes, but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money? The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies. Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway. Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know, again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there there&#039;s value somewhere in here. And we&#039;re being asked exactly the question, where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be advertising instead of pornography. It surprised a lot of us. Um, you know, some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene. haven&#039;t been there, you should check it out maybe of it. semantic technologies. It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question? Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data. The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because they one person saw in the data and it turned out to be something really of interest. I think actually we can point to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and then when he decided off through his own intuition that this would change the way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
&lt;br /&gt;
I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6646</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6646"/>
		<updated>2026-04-17T13:50:55Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
&lt;br /&gt;
Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
&lt;br /&gt;
It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take.&lt;br /&gt;
&lt;br /&gt;
So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know, he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms. Ontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6645</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6645"/>
		<updated>2026-04-17T13:48:36Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
 So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not&lt;br /&gt;
 showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert&lt;br /&gt;
 themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown&lt;br /&gt;
 said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
 Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the&lt;br /&gt;
 site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
 It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened&lt;br /&gt;
 here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m&lt;br /&gt;
 not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by&lt;br /&gt;
 the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of&lt;br /&gt;
 tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three&lt;br /&gt;
 months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil&lt;br /&gt;
 servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
 We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know,&lt;br /&gt;
 he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms.&lt;br /&gt;
Anontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6644</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6644"/>
		<updated>2026-04-17T13:48:17Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the&lt;br /&gt;
 updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York, they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
 So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not&lt;br /&gt;
 showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert&lt;br /&gt;
 themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown&lt;br /&gt;
 said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
 Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the&lt;br /&gt;
 site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
 It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened&lt;br /&gt;
 here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m&lt;br /&gt;
 not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by&lt;br /&gt;
 the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of&lt;br /&gt;
 tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three&lt;br /&gt;
 months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil&lt;br /&gt;
 servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
 We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know,&lt;br /&gt;
 he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms.&lt;br /&gt;
Anontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6643</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6643"/>
		<updated>2026-04-17T13:47:57Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that&lt;br /&gt;
 have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the&lt;br /&gt;
 updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York,&lt;br /&gt;
 they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
Evan Sandhaus:&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
 So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not&lt;br /&gt;
 showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert&lt;br /&gt;
 themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown&lt;br /&gt;
 said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
 Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the&lt;br /&gt;
 site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
 It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened&lt;br /&gt;
 here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m&lt;br /&gt;
 not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by&lt;br /&gt;
 the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of&lt;br /&gt;
 tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three&lt;br /&gt;
 months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil&lt;br /&gt;
 servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
 We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know,&lt;br /&gt;
 he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms.&lt;br /&gt;
Anontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6642</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6642"/>
		<updated>2026-04-17T13:47:36Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled. And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that&lt;br /&gt;
 have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the&lt;br /&gt;
 updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York,&lt;br /&gt;
 they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only&lt;br /&gt;
 problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
 Evan Sandhaus&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
 So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not&lt;br /&gt;
 showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert&lt;br /&gt;
 themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown&lt;br /&gt;
 said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
 Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the&lt;br /&gt;
 site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
 It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened&lt;br /&gt;
 here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m&lt;br /&gt;
 not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by&lt;br /&gt;
 the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of&lt;br /&gt;
 tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three&lt;br /&gt;
 months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil&lt;br /&gt;
 servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
 We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know,&lt;br /&gt;
 he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms.&lt;br /&gt;
Anontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6641</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6641"/>
		<updated>2026-04-17T13:47:19Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled.&lt;br /&gt;
 And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that&lt;br /&gt;
 have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the&lt;br /&gt;
 updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York,&lt;br /&gt;
 they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only&lt;br /&gt;
 problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
 Evan Sandhaus&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&amp;lt;br&amp;gt;&lt;br /&gt;
 So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not&lt;br /&gt;
 showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert&lt;br /&gt;
 themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown&lt;br /&gt;
 said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
 Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the&lt;br /&gt;
 site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
 It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened&lt;br /&gt;
 here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m&lt;br /&gt;
 not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by&lt;br /&gt;
 the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of&lt;br /&gt;
 tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three&lt;br /&gt;
 months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil&lt;br /&gt;
 servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
 We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know,&lt;br /&gt;
 he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms.&lt;br /&gt;
Anontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6640</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6640"/>
		<updated>2026-04-17T13:46:48Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Transcript of lotico session [[Data Gov - Bringing Government and Scientific Data to the Web]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled.&lt;br /&gt;
 And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that&lt;br /&gt;
 have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the&lt;br /&gt;
 updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York,&lt;br /&gt;
 they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only&lt;br /&gt;
 problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
 Evan Sandhaus&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann&lt;br /&gt;
 So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not&lt;br /&gt;
 showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert&lt;br /&gt;
 themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown&lt;br /&gt;
 said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
 Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the&lt;br /&gt;
 site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
 It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened&lt;br /&gt;
 here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m&lt;br /&gt;
 not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by&lt;br /&gt;
 the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of&lt;br /&gt;
 tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three&lt;br /&gt;
 months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil&lt;br /&gt;
 servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
 We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know,&lt;br /&gt;
 he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms.&lt;br /&gt;
Anontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6639</id>
		<title>Data Gov Transcript</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_Transcript&amp;diff=6639"/>
		<updated>2026-04-17T13:45:57Z</updated>

		<summary type="html">&lt;p&gt;Marco: Created page with &amp;quot;  Gale A. Brewer:  Thank you very much. This is a real honor to be here and I&amp;#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&amp;#039;s going on in city government particularly from the legislature side. I&amp;#039;ve been in the city council since 2002 and I&amp;#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&amp;#039;s almost an ox...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer:&lt;br /&gt;
&lt;br /&gt;
Thank you very much. This is a real honor to be here and I&#039;m certainly not going to be half as knowledgeable as the speakers who come after me, but I will try to give you a picture of what&#039;s going on in city government particularly from the legislature side. I&#039;ve been in the city council since 2002 and I&#039;ve always worked in government or a little bit in the private sector but always believed that government data should be public and that&#039;s almost an oxymoron because so much of it is not. So even years ago I built the first city website from scratch and even put up some little preliminary kiosks years ago in front of city hall. it had to be temporary because otherwise the landmarks commission would get upset. So you always have to work around things.&lt;br /&gt;
&lt;br /&gt;
But I&#039;ve always believed in this idea of public information should be public and when I came into the council I was head of made head of the u at that point it was a subcommittee on technology and then for the next eight years it was one of the 40 city council committees. There are actually 51 members of the city council for those of you who don&#039;t have the pleasure of spending all your time in city hall. there are 51 members and the budget of the city now is now 63 billion. There are about 303,000 employees. It&#039;s the fourth largest budget in the United States. United States budget is bigger. State of California, state of New York, and then New York City. So it&#039;s bigger than Massachusetts and Cal and Texas and Florida and so on. So it&#039;s really like a and half the countries of the world were bigger than half the countries of the world. So technology plays a role. Needless to say, I think government is always slower. You know that from those of you who either work for government or work with government or try to get information. So for the last eight years, we&#039;ve been holding hearings when really nobody had ever had hearings on technology in government before. We had hearings right after 911 about how technology helped as much as was possible in that horrific situation in the aftermath tracking and Verizon got back up pretty quickly and all the things that 9/11 had to deal with and all we&#039;ve had hearings on of course the maps and we&#039;ve had hearings on clouds and we&#039;ve had hearings on spectrum and we&#039;ve had hearings on the everything the franchises with Verizon and the cables. I mean there&#039;s no topic we haven&#039;t kind of looked at but the one that and broadband we brought passed a bill that said that we had to have hearings in all five buraus to see whether there was or was not accessible lowcost broadband in the five buraus and of course there&#039;s lots of it but it&#039;s expensive and not very fast and so it&#039;s sort of like a a a mish mash of different topics, but the one that is always hiding I think a little bit is the data. there is in the city anybody here work for the city of New York? Are there people besides So there&#039;s a couple of us couple of us. the city of New York does have a very vibrant I guess it&#039;s a closed meeting but I go fairly regularly and it&#039;s all of the CIOS in the different agencies who meet on a regular basis and it really amazing I think because these are the men and women who have to make things work for this huge city and you know they&#039;re they&#039;re under represented in the you know newspapers on the positive note every day. But there I don&#039;t know how many  we have 80 city agencies probably more but the official ones are 80 and so there&#039;s probably there&#039;s obviously many people and this group comes together I think it&#039;s called the municipal data council and they come together with the commissioner of do it and other agencies on a regular basis to talk about what their issues are and what their challenges are and how they can work together and how things aren&#039;t working and so on and there&#039;s a big push now to try to make more centralized because the agencies police departments the worst but don&#039;t tell anybody I said that they&#039;re very siloed you know they want to do everything in their own little world and I&#039;m sure that&#039;s true in academia but it&#039;s really true with city agencies so given that whole backdrop and based on other cities we&#039;re always looking New York City thinks we know everything but we do try to think particularly on the tech front are there other cities that are doing interesting things and I think on the issue of open data No, that&#039;s true. Now, the mayor, uh, Bloomberg did, as you know, a couple of times, at least once, and I think there&#039;s another one coming up, and some of you may have participated, um, you know, had an apps conference. I went to the opening, uh, of one, and it was it was exciting. The only problem for me, and I&#039;m just going to speak for myself, was it was all based on parking spots and I don&#039;t know, God, things that, you know, they&#039;re important, but I&#039;m interested in how you can help New Yorkers. Maybe getting the parking as part of parking. But anyway, the issue was the data that was available was only those that either the city thought was easily lowhanging fruit or one that u somebody who was doing the apps had requested. Now if you don&#039;t it&#039;s like anything else if you don&#039;t know the data is there how are you going to request it? So it was it&#039;s a very small subset of all the data that&#039;s available. So we introduced a bill that said that the data has to be available.&lt;br /&gt;
&lt;br /&gt;
We did it in 2009 and we did it again when we started a new session in 2010. So now we&#039;re actually had a hearing in June this of this past June and the issue is whether we can pass intro 29 of 2010. And what this legislation says is that the city of New York needs to make its data available in a format that is accessible to the general public, not in some formula that nobody can understand, even with the apps. And some of you may have won the contest or participated in the contest. You had to be one of you in the room to understand how to make it something that the public could understand. So we I obviously my background is human services and housing and things like that. So I want something that people working in the nonprofit sector can use to figure out where&#039;s the affordable housing how many homeless do we really have and what are their needs and so on and so forth. So at this moment just to give you an update and I can go through the bill in a minute but at this moment we&#039;ve had a couple of meetings with the new commissioner Carol Post. she came from the office of operations. She&#039;s been there, I think, since December of last year, and she did say in her opening statement to the city council that she wants this bill to pass. Now, of course, you always hear that and you want to make sure that it&#039;s this bill or something that&#039;s much too narrow, in which case it&#039;s not this bill, but to her credit, she has had meeting with us, meaning the city council, and she has met with all of the so maybe 80 or fewer agencies, but all the agencies, and she&#039;s giving them time, a little bit of time, not much, this fall to meet goals or to come up with goals. so they can have open data. And the idea there would be she&#039;s got certain criteria that they have to meet and the agencies have deadlines. and when we&#039;re going to meet again, I think it&#039;s either the end of September or the beginning of October to figure out uh which agencies are meeting their goals and their timelines and which are not. Now, there are some agencies, and you probably know this, when you call 311, has anybody ever called 311? You probably Okay, good. Sometimes it works and sometimes it doesn&#039;t. But it is generally u an amazing talk about databases that is an amazing database. but those sometimes the agency can quickly respond and sometimes it&#039;s a legacy system and it has to go through a whole bunch of channels before a you get an answer but b you get data. So she&#039;s dealing with some agencies that have some very old hardware and is not compatible with anything that&#039;s helpful. So she&#039;s trying to figure out how to make this data accessible in the broadest format. And of course we&#039;re pushing very very hard. Secondly, costsaving. Anything that saves money in today&#039;s world is a good thing. If you haven&#039;t foiled freedom of information law, you may be the only New Yorkers who&#039;ve never foiled because many, many New Yorkers file. Reporters file, journalists file, upset New Yorkers file. Um, I used to work in a city agency that was always getting foiled.&lt;br /&gt;
 And that is a very timeconsuming response because often it&#039;s given to be honest with you to a low-level intern.&lt;br /&gt;
&lt;br /&gt;
And that intern has to gather all the material that the upset person or journalist is interested in and supposedly gets back to you in a certain time period, but you can always get an extension. So it&#039;s my just one example and it&#039;s very timeconuming because you got 80 city agencies often corporation council has to get involved if it&#039;s a more high level request and of course what is responded to and what&#039;s blacked out is obviouslyimportant. Um, so the fact of the matter is that I think that if you have a lot of this data up on the web in a format that people can understand that you will not get as many foils and you won&#039;t have to be answering all of these endless and sometimes important and sometimes there&#039;ll still be foils, but it&#039;s a cost-saving measure and I want to add that because people those of us who just want the data out there to be able to work on it and use it practically and to help save lives and make people&#039;s lives earlier here don&#039;t see it also as a cost-saving measure. So I think that&#039;s extremely extremely important. the other thing I want to mention is that there is you know there&#039;s a lot of interest in this legislation. Obviously people in the public sector who are writing about what the city is or isn&#039;t doing or interested in it. people from like Common Cause and the groups that work on Citizens Union and groups that work on public access have testified over and over in support. and I think what comes up again and again is what kind of format that the data has to be in. And that&#039;s something that some of you in the room might have some good ideas about and certainly something that the city is trying to take note of. That is definitely part of intro 29 of 2010 is having data that&#039;s in an accessible format. the other issue I think is you know how do you how does the city decide and how do we advocate to know which agencies will produce and which not it&#039;s always challenging with the police department. and I obviously were concerned about health records. Something that just sort of an aside some of you may have seen checkbook NYC. It was just put up a couple July 1st by controller John Louu and Checkbook NYC is an amazing amount of data. Not something that you need to manipulate at all, but it is every penny that the city spends every day. It&#039;s updated daily. It&#039;s about 35 billion that&#039;s up there now out of a budget of 35 billion. But this is about 35 billion. And it has you can search, you know, it says that the department of health has spent money on pharmaceuticals and doctors and pediatricians and that the mayor&#039;s office has bought liquor, but they&#039;re going to get reimbured for a party. And it says how many car services the board of elections has used when the workers go home late, etc., etc. For those of us who interested in gossip and, you know, things like that, it&#039;s fabulous. So for those of you who are interested in more mundane things like tech and things that a little bit more professional, it&#039;s also there. what&#039;s not there because they&#039;re working at it is the salaries of the workers in a because there&#039;s a lot of concern about that kind of privatization and they&#039;re worried about whether people would be stocked or home at home home names if that kind of thing is still not up there and it will be. But I all I I mention that first of all it&#039;s a fun site. You should look at it. But I also mention that because I think the city is finally getting the message that data that&#039;s public needs to be public. And so one part of it is where does the city spend their money?&lt;br /&gt;
&lt;br /&gt;
Where does your tax money go? And already just the other day somebody did a story about car washing because it turns out the Daily News must have seen a police car being washed. I don&#039;t know how the reporter started on the story, but then he realized that there are like 10 different agencies and it&#039;s a hodgepodge of which government car is washed at which car wash. And then unfortunately somebody got a car wash for like $126 or something that&#039;s very expensive car wash. And so from there it was a story, but he got all of his information from putting the word car wash or some something like that into checkbook NYC. Um, so that&#039;s good for the public to know and hopefully we&#039;ll have some hearings and maybe there&#039;ll be vendor car washing in terms of the value for the dollar. But this particular legislation is is larger in the sense that it is looking at the data sets that you as people who know how to academics and your jobs figure out how you can take this data and make it useful to the public and that will never get done by government. Um, and so it&#039;s absolutely necessary. It&#039;s your data. It&#039;s your right to have that data. But at the same time, in this in today&#039;s world, what I think is so exciting about it, it can create jobs and it can create information and hopefully make people&#039;s lives better. So on so many different levels, this is a very exciting opportunity. how we get to that point is what we&#039;re advocating for as much as possible. and then just finally, you know, there&#039;s so many questions that&lt;br /&gt;
 have to be answered. What&#039;s the record policy? How long do you keep the data for? Um, we&#039;re already running into some of those questions even without a large large amount available to the public. What are the technical standards? Um, is it based on agencies or is it across the board in terms of the standards? Is it one data portal, XML, raw data, support, RSS? Um, and you know, just something that&#039;s readable. And then of course I would be a big believer particularly in the nonprofit sector, you&#039;d actually have to have some classes and some training. I think for some of the nonprofit sector people who could really use this data to be able to use it more effectively. I think that the Washington DC apps for democracy, some of you may know that has done has actually saved money and is doing the kind of work that we would like to see done in New York. A lot of times the&lt;br /&gt;
 updating is a problem. If you ever go to a government website, you will find at least in some cases, it&#039;s not updated as often as some of the ones in the private sector because there&#039;s obviously a perhaps a bigger motive to try to get the private sector ones updated. I find I always tell my staff, don&#039;t rely just on the web. You got to pick up the phone to get the real data because it&#039;s not necessarily going to be correct. We had a big challenge even with the 311 data. Um, the community boards. Has anybody ever been to a community board meeting? I hope somebody like three people have been to Oh, four have been to a community board meeting. I I&#039; I&#039;ve probably been to 5,000 community board meetings. So, just to give you an example of the challenging of of elected office, but there are 59 community boards in New York City. And for those of you not from New York,&lt;br /&gt;
 they&#039;re like little little city halls.&lt;br /&gt;
&lt;br /&gt;
There&#039;s a staff of about three people and then there are 50 citizens. Anybody who is interested, you can apply to the B president and get on and if you&#039;re been to some meetings and you have an interest in putting some time into the neighborhood, you&#039;ll learn a lot. But there&#039;s a lot of data there. Um, you know, the street lights go out, the road needs paving, um, the applications for zoning, the applications for enclosed and unenced sidewalk cafes, and I could go on and on. at the same time and those often are either dealt with locally, passed on locally, but certainly have a knowledge locally. You even have situations where the movie people come in and if you&#039;re on the west side, everybody hates them even though they&#039;re great for the economy. How many of those permits are let on any given moment in Manhattan? Because we always want them to go to the Bronx or Brooklyn. Please go somewhere else. But you know, that kind of information would be so helpful to regular New Yorkers. why am I getting all the movies and I can&#039;t get a parking space? You know, that&#039;s the kind of thing that people really actually want to know. but the community boards are not in real time with 311. So you call 311 and the community board doesn&#039;t know it. So, we&#039;re trying to think of ways that community boards which are run by city employees can in fact have that real-time data because often this is an example of the importance of this data sharing because the community boards would like to know have a 100 people  called about the street light or am I the only am I only getting one complaint or is are 100 people complaining about a road that&#039;s not paved or am I the only is the only complaint coming from me because if there are of people calling then we&#039;d like to be more active in trying to solve the problem. So there&#039;s also something called scout. These people in the city I&#039;ve never actually seen them. They run around little golf carts and they find problems but they don&#039;t share them in real time with the community boards. They give the data in centrally. So I&#039;m giving you some examples of some of this huge data opportunity that&#039;s out there and that we&#039;d like to capture with intro 29 in 2010. So, I I&#039;m going to stop there. We can certainly talk about it later. I would love if people have the time to either write to the mayor or to Speaker Quinn to say that you&#039;d like to have intro 29 of 2010, you can send an email passed in some form because you think that having government data be public is a good thing. Thank you very much. All right. Thank you so much for coming. I just want to make sure that you want to take one or two questions if you have one. I think it&#039;s interesting that you characterize the police department as being behind this sort of stuff because the public perception is that internally like blockby crime data and predictive analysis they&#039;re actually cutting edge on this. Yeah, I I wasn&#039;t correct. They&#039;re not behind. They won&#039;t share it. That&#039;s the problem. In other words, they&#039;re very the comat is excellent. obviously they&#039;re doing a great job on terrorism. So now they have 1,000 police officers as we&#039;re sitting here on computers. So that could be everything from comat to looking at porno and  dealing with those crazy people. Porno meaning not them but making sure that others are not doing it. Cutting and also looking at the issue of terrorism 1,000 officers. That&#039;s a lot of people. So, no, they&#039;re very and they have NYC win, which of course is the system that I didn&#039;t talk about, which is a citywide a first responder system wireless, and they&#039;re able now in the near future to have much more connectivity and instant real-time info in the cars. But they&#039;re the only&lt;br /&gt;
 problem is they don&#039;t like to share anything. That&#039;s the little problem.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann:&lt;br /&gt;
&lt;br /&gt;
All right. So, with that, we are heading over to Jim. Thank you so much for coming. Thank you. There&#039;s so much work to do. So, um, when is the next meeting? I will let you know then. I don&#039;t know if it&#039;s been set because we&#039;re starting a new session during the fall, but we&#039;ll let you know and then you can circle. Yes, we can. So, we&#039;re switching the screen. So, we have a maybe Evan Sandhaus, you want to say a few words about September 30th or you want me to do that now, right? I&#039;ll come second.&lt;br /&gt;
 &lt;br /&gt;
 Evan Sandhaus&lt;br /&gt;
 &lt;br /&gt;
I&#039;m Evan Sandhaus. I am the I&#039;m I&#039;m one of three assistant organizers of the semantic web meetup along with Marco and I am we&#039;re going to do a pretty cool event this coming this the first public announcement of it the meetup page isn&#039;t out there yet but on September 30th we&#039;re going to do a session in jointly with a new meetup called Hacks and Hackers. I don&#039;t know if any of you have heard of them or not. They&#039;re a they&#039;re a new meet up focused on the intersection of journalism and technology. the journalists are the hacks. We&#039;re the hackers. so it&#039;s called hacks and hackers. and we&#039;re going to do a joint session with them on the role of data and metadata in the collection and in the production and management of news. So we&#039;re going to have we&#039;re going to have speakers from right now could change but right now we have the New York Times, the Wall Street Journal, Hurst and AOL News lined up to speak and it&#039;s going to be a really really cool event. And I pretty sure we have a sponsor that&#039;s going to pay for drinks, too. So, it&#039;s going to be it&#039;s going to I if for no other reason than that, I I I hope you guys feel like this is going to be interesting. the announcement will go up on Meetup shortly as soon as I can get around to writing the description of the event. So, I hope to see you all there and should be good. Where is this going to be? All right. So, it&#039;s going to be at AOL&#039;s at at an AOL facility. I don&#039;t know exact. It&#039;s here in New York. We&#039;ll hear more about this.&lt;br /&gt;
&lt;br /&gt;
Marco Neumann&lt;br /&gt;
 So again, so welcome, Jim.&lt;br /&gt;
&lt;br /&gt;
Jim Hendler:&lt;br /&gt;
&lt;br /&gt;
So I&#039;m going to kind of wander through a few sites and things like that. So to to duplicate my talk on your own, the only thing you have to remember is data.gov, which is the big letters up there. And that&#039;s where I&#039;m going to start. Um, if you look in the corner, most of you can read it up here where it says an official website of the United States government. So, I&#039;m not showing you my stuff now. Well, I am, but I&#039;m not&lt;br /&gt;
 showing you primarily my stuff. This is a site run by the GSA out of the White House, out of OSP in part, out of several other agencies. The first act taken by Obama was to create a new CI  was to hire the first ever US CIO and CTO. The US CIO started this idea that within the government there would be the release of data sort of like you just heard about from New York. the CTO is is  involved in an even larger open data government aspect to also go beyond data to a lot of the documents a lot of the  policies processes etc. So there&#039;s a huge amount of activity going on trying to figure out how to give people stuff and particularly the data set. Now this is the historically what happened is so Obama created this in May 21st.&lt;br /&gt;
&lt;br /&gt;
This is 2010 now 2009. So a year ago May the Obama administration announced the creation of the data gov website with about 50 data sets on it. So about 50 government data sets have been made available and they were made available in various formats primarily though just raw data raw commaepparated variables some of them XML things like that. Over the next few months another I forget the exact a couple hundred more came on. Now, right about that time, the Obama administration  issued what was called the open government directive, which said that every agency in the executive branch of the government, which is the great bulk of the agencies you&#039;ve heard of, DoD, DOI, DO, bureau of agency of so and so, all those guys had to create find basically it it  ramps in and it&#039;s a long story, so I don&#039;t have time to go through the whole thing, But essentially they had to identify some highv value data sets and there&#039;s a definition of highv value they had to release those data sets through  data go so by roughly December January time frame there were about 1100 or 1200 data sets on data.gov now meanwhile over in Britain data.gov.uk UK was forming and data UK.&lt;br /&gt;
&lt;br /&gt;
You can see data. UK if I do this data. UK. oh boy. These speeds are going to make this demo fun. So this is all live by the way. Nothing nothing on my sleeves. I should go to my PowerPoint with these speeds. you will notice if if a Whoops. See if I touch the side. Let&#039;s try. That&#039;s the advantage of having power, but it&#039;s not nearly as effective. Okay.  Data go UK, official site of Her Majesty&#039;s government. you may recognize the RDF logo on it down there.  so what happened is s let&#039;s see May, June, July. So so late summer of 09 Tim Berners-Lee was at a meeting with Gordon Brown, the prime minister of England and some other relatively well-known people. Bill Clinton was there. I&#039;m told I wasn&#039;t there. Um, and they were kind of going around talking about things Britain could do to kind of get attention back and reassert&lt;br /&gt;
 themselves as a world leader and all that stuff. And Tim said what he always says to governments, you know, release all your data. He said Gordon Brown&lt;br /&gt;
 said, &amp;quot;Great idea. Let&#039;s do it. Come back and tell me how.&amp;quot; Tim said it was the scariest moment of his life. No one had ever said yes before. And and part of what motivated, of course, was that the US was doing it. So there&#039;s been kind of what Tim refers to as the friendly rivalry, what I call the war of 2010, but the the it&#039;s it&#039;s kind of cut and fun because for those of us who&#039;ve been helping the governments, so at in Britain, the data sets are released from the beginning, many of them in RDF. Okay, they all must be in a machine readable format, but they&#039;ve mandated RDF as the primary machine format. a semantic web from the get-go.&lt;br /&gt;
 Now, over on the US side, we were not mandating anything like that or doing anything along those lines. But what happened is is the laboratory at RPI that the three of us speaking tonight represent I moved to RPI from the University of Maryland where I was for about 20 years in 2007. We started what&#039;s called the Tetherless World Constellation. I&#039;ll say a couple minutes about that in a second. Deborah McGinness and Peter joined us over the next couple years. Dre Luciano has just joined us. So that&#039;s sort of the faculty of this center. We have about 20 30 grad students now depending exactly what you count staff faculty and what happened was we started looking at these data sets and said well why don&#039;t we do what the bricks are doing and start turning them to RDF and showing what these guys showing these guys what we can do. So my lab, we just sort of we didn&#039;t have any funding. We didn&#039;t have any we just, you know, had been saying for a decade now that if people would just release the data, we&#039;d be able to mash it up. And we thought, well, maybe we&#039;d see if what we were saying was really true and was this great great revelation to me that actually all this stuff we&#039;ve been saying about the semantic web actually works. Uh, you know, there&#039;s more of you in the room than the first semantic web meeting I ever held. I mean, you know, it&#039;s really exciting. Um, that was an international semantic. Anyway, so what happened was for May 21st, 2010, the first anniversary of data gov, there was a relaunch of the&lt;br /&gt;
 site and the the CT the CIO of the Vet Kundra was looking for applications.&lt;br /&gt;
&lt;br /&gt;
So he was looking for where has some interesting stuff happened and there had been a little data gov newsletter every week something go on and he started noticing that about half the issues of the data gov newsletter talked about some demo that had built at rpi been built at rpi so he asked me to come meet with him which was exciting I don&#039;t get called to the white house too often press nicer and and the upshot is we showed him what we were doing we had built over the course of about six months months using primarily undergraduates and graduate students who had never touched the semantic web before they joined the lab. 40some demos that mash up significant amounts of US government data. We also had built a converter to start converting data sets to RDF and had created about 6.4 billion triples out of the first few hundred first about thousand data sets we had converted. So I got asked to come in sort of more officially. I&#039;m now a I&#039;m not a let&#039;s see I&#039;m not an adviser and I&#039;m not a consultant because those two words have meaning. I&#039;m an expert. That&#039;s what so I am the internet web expert for the data gov project. And what I&#039;ve been doing is helping them look at some of this. Now meanwhile they also wanted a number that would beat the pants off the bridge. So we now have 272,000 data sets available for you on this site. Now it&#039;s worth noting that they come in three forms. So there&#039;s sort of raw data, tool catalog, and geo data. About 270,000 of the 272,000 data sets there are geo data. At the moment, they&#039;re not very well organized.&lt;br /&gt;
&lt;br /&gt;
Searching for anything is very difficult. So a lot of what the the priorities are at data gov now is how do we make that stuff as much fun and as easy to play with as we did with the first set of data. But let me show you what was going on. So, um, so you can find your way to the group at RPI. We&#039;re called the Tetherless World Constellation. I won&#039;t go into that too much. So, now we&#039;re leaving the government site. Now, we&#039;re into the site at RPI that Deborah Peter and I run. Um, so this is our lab. There&#039;s various inundry there you can see. And what we&#039;re working on is sort of two things. One is of course a lot of semantic web. All three of us are known for semantic web technology and we continue that axis. But we&#039;re also looking at spawning out towards some other things. You&#039;ll hear about a couple of them today. What could we do when we start doing a lot of data stuff? what is the whole science of the web? I won&#039;t go into that one or I&#039;ll be here for another hour and these guys will get mad at me for not letting them speak. And supporting science. So I like to say you know what I&#039;ve been looking at is very very broad data integration. What Peter will talk about later is more a small number of data sets whose size make make the entire government data release look tiny and Deb is sort of somewhere in the middle in terms of the technologies we need to really do these so the ontologies the provenence things like that so that&#039;s sort of tonight&#039;s talk so I&#039;ve sort of segueed from halfway through my talk back to the introduction now back to my talk from the data gov so from our website which you can get to from data gov. You&#039;ll see a pointer to the data gov project or you can just remember data-gov.cw.rpi.edu and we we we heard what you were asking for Gail before you even said you needed it. So what we&#039;ve been doing on this site is building mashups of government data. We started by first just building visualizations putting them to RDF figuring out what we could do. But the goal of course of RDF is the data integration and the data linking and and all that kind of data stuff beyond you know sort of looking at one database through one data set through one data portal. So you can see we&#039;ve currently converted about 687 of the 2,769 data sets that are in that are not in the geo data sets. So the geo data have its own formats and things like that. We&#039;re looking now at how we&#039;re going to bring that stuff to the semantic web. Created about 6 6.39 billion triples.&lt;br /&gt;
&lt;br /&gt;
Now those triples are also now available from data gov. I won&#039;t go back there and show you, but if you go down to the semantic web pane at data gov that you know so so we built a bunch of demos right if you come here you can see our demos but being an academic organization we actually want to do more than just build demos which is we want to make it so people could build their own demos. problem is I can&#039;t show you all of our demos tonight because this thing only has IE IE isn&#039;t friendly to Google visualizations. There&#039;s a big fight as you know between Apple and Flash. Less known is the fight between Microsoft and Google. And it all depends on, you know, sort of what you&#039;re going to be able to see for what kind of phone when. But but if you got Firefox or Safari and you go here, you&#039;ll see some some stuff in these demos that you&#039;ll see partial visualizations here. So occasionally I&#039;ll click on this and nothing will happen. I don&#039;t know which ones are friendly or not, but but this is a good example. This is actually our best known example today for an interesting reason I&#039;ll show tell you about in a second. So what you&#039;re seeing here is ozone sensors. So the the government released a data set of for a whole bunch of sensors what were the ozone levels being reported by the sensor. The size of the circle is the average based on the average value. So the glance you can see this. Now the interesting thing about this data set is it didn&#039;t say where the sensors were right it just had the sensor. So this data set just had the values and the external key was a sensor. Okay. We did a few web searches.&lt;br /&gt;
&lt;br /&gt;
We found there was another EPA site that actually had the had the locations. So obviously mashing up the sensor values with the sensor locations let put them on the map. We also know a lot about the terrain in the world and things like that because that data set is also available elsewhere. So what you have on the side here is a a faceted browser of those data sets. so here are the ones that are on mountain tops, things like that. But some of them were from EP these are the ones that the EPA maintains. These are the ones that the National Park Service maintains, things like that. So we had to get some stuff from an EPA website, some stuff from a park service website. If you go into one of these, what we can do being it&#039;s, you know, all web stuff, this will let us actually is terrible. so what I&#039;m doing now is I&#039;m doing a sparkle query against an endpoint. This one is actually running since I launched it from our site on our endpoint. If you go to the data gov site and click on semantic, you&#039;ll see about seven of these 43 demos. they&#039;re running on their own virtuoso server. So they&#039;re using their own triple store locally for the government. Um, and they also host, as I said, that you can get gzipped files of all of the triples that we&#039;ve converted. Uh, Kingsley Idaho, for those of you who know him, those of you who are insiders in and this stuff know him. Uh, he hosts them all live and also host the links between them and the rest of the, uh, the link data linked open data cloud. So you about 13 billion triples of linked data out there. Six billion are from this project so far.&lt;br /&gt;
&lt;br /&gt;
When we start going after this 270,000 data sets I don&#039;t even want to think how many triples it&#039;s going to the metadata they have very skimpy metadata we estimated just to convert the metadata would give us another few hundred million triples okay my guess is what we&#039;re seeing here is an IE incompatibility so what I&#039;ll do is I&#039;ll go back to the IE friendly version back at the real datab show you a couple other demos. So again, if you click on the semantic web tab, oh, so let me mention for a second community these are the organizations working most closely with data. Open gov is is a larger government initiative or RPI, the Sunlight Foundation, the World Bank has also done a major data release. One of the things we&#039;re working on now is hooking up some of their things. These are the US states that have published data out there. You can see them on my map somewhere. I guess you can&#039;t see states. These are the countries that have done data release. So far, the US and the UK are the only two that have native where you can actually get it in semantic web formats natively. However, for all the others, if you want them, the converter we built, which is available through our website, is available for you to take whatever you want and play with. Let me just show you what some of these demos look like. So again, I was showing you CastNet was this one. See if their site will actually work better. So what happened is when we ported our stuff over their site because so many government users have IE only. Where&#039;s reload?&lt;br /&gt;
 It&#039;s next to the So that one that one&#039;s odd. That one is actually a Google error. The problem there is coming as best we can tell from the data set. F5 good. Okay. So this is so what happened&lt;br /&gt;
 here is we went off to a different and EPA site. So of course now suddenly the data can be linked to other web stuff. It&#039;s URIs of URIs. You can see it. So the mashup. So what we have now is we&#039;ve take So if you&#039;re counting, there&#039;s the EPA site, the National Park Service site, the EPA data set, that&#039;s the primary sensor values. Now we&#039;re off to EPA and NPS sites themselves. So we can link out to them. And the other thing we can do is get down to the raw data readings. Raw data readings. Again, we&#039;re we&#039;re querying. These are these are pretty big on the the internet&#039;s slow here, but this usually takes about three four seconds. So, I&#039;m&lt;br /&gt;
 not sure what&#039;s going on. Uh, again, I don&#039;t know what&#039;s exactly I friendly, what&#039;s not, but so what&#039;s what you should be seeing here is a nice timeline slide. I&#039;ll show you a different. Okay. Okay. So, some of the other demos that they&#039;ve picked up is so here&#039;s this just state library books by state. If we click on one now, you can&#039;t see the graphs. It would launch a graph that would show you some information about it. I was really bad. I&#039;m I&#039;m sorry. I mean that is not an official policy of the US government. they like I so here you know so here what what we have here is this is broadband adaptation in urban versus rural areas for the different states there&#039;s a color coding here that you can sort of see on the map we go to here&#039;s here&#039;s a one this is global foreign aid by the US and if you go to our site you can actually see the actually this one&#039;s probably dispatching to our site yeah it is Um, Peter, mind me to kill D when we get home? Uh, I don&#039;t know what&#039;s happening. Someone&#039;s playing around with our our end point, I think. Oh, here we go. All right. So, um, there&#039;s a map here somewhere for this. This one actually is an IE problem, but um, okay. So for example if we want to see what the aid to what&#039;s good India is good so India has been funded US aid this is what&#039;s been done by&lt;br /&gt;
 the department of agriculture department of state their categorizations now one of the things that&#039;s really interesting about this stuff is that the visualizations I&#039;ve been showing you are all just standard APIs that are supported by Google some of them are the exhibit API from MIT we&#039;ve also gone out to Yahoo types, Yahoo Boss. &lt;br /&gt;
 &lt;br /&gt;
So lots and lots of people now are building sites where if you give them an XML page, they will draw it to to to one of these as long as you just format things right. And basically to get to the formatting right, we we&#039;ve built a bunch of stylesheets. All of those are available. So what I was trying to say before is if I go back to our website primary site and I&#039;ll finish with this You can see we also have for example a bunch of&lt;br /&gt;
 tutorials. So if you want to learn how to do all the stuff I&#039;ve just shown you, granted where most of these are written by graduate students, undergraduates, me, you know, a lot of illiterate people. so they&#039;re not, you know, quite in the format you&#039;d get if you went to a professional site. But how to build these things? if you want to know you know how our endpoints work and things like that this is some of the things about gov these are some of the external things we&#039;ll make available to you if you&#039;d like to see some videos about creating the website you know here&#039;s this is a undergraduate from Bennington who came to our lab for three&lt;br /&gt;
 months is a political scientist never touched computer science before built a whole lot of really cool things working with my things maps with Twitter feeds over them and things like that so You could see broadband versus so we had what the the broadband adaptation adapt broadband adoption in various states in the US versus how many Twitters were coming from those various states things like using their geo stuff. I mean all this stuff is easy. That&#039;s the point I&#039;m trying to make. In fact last week we we held the first US government mashathon. They didn&#039;t like the term hackathon. Where we had is is 30 30 civil servants mo half civil&lt;br /&gt;
 servants half support people for them who came in. We taught them in in a day and a half how to build their own demos. And I you know I should have loaded it up. I&#039;d have to do a little bit of navigation to find it. But we have now have five or six new demos that were built by government people in a two-day period. They did need a little tutoring. but again we&#039;re we&#039;re we&#039;re not at rocket science anymore. The semantic web stuff has reached a point of maturity. You can get your hands on it. we have a lot of stuff about how to do it. You can cut and paste. So in in the demos every demo has if you go to it oh here it&#039;s evidence here. I don&#039;t know if this this one actually won&#039;t show up with live because of the IE issue. What we can do is we&#039;re using some of the New York Times API. So here what you have is the budget of an agency and here what you have is New York Times stories about those agency budgets and now we&#039;re working on New York Times link data to make it so for people and things like that we can get the reput you know the so this is just a keyword search there&#039;s now a better way to interact with their stuff. So again, so if you want to know why in this particular time, I don&#039;t know what agency this is, the American Battle Monuments Commission, right? What were they doing? Well, this was something about refurbishing the Vietnam Memorial. So you can see where that money was being spent. Um, for each of these demos, what you can see is we have what technology we use. Here&#039;s the actual sparkle query, right? So if you want to cut and paste that against our endpoint or against another endpoint, you can. Lots of semantic data there. So again, so I invite you to come play. Lots of government data. we&#039;re getting better and better at this. Our our goal in this project now is to make it so developing one of these government mashups is roughly the same level of complexity as creating a website, right? We&#039;d like it to be as turnkey as that. I mean, you know, so it&#039;s not necessarily that every end user can do it, but certainly every web master should be able to do this. Certainly everyone in an organization who has enough literacy to understand the underlying, you know, so if you&#039;re someone who who looks at a web page and can&#039;t tell that there doesn&#039;t know what the what that there&#039;s some kind of machine readable format down there that presents it to you, you know, you&#039;ll need a little education. But we we really think we can get there. We&#039;reware where from scratch on your own playing with our stuff. Those of you who have any familiarity with this stuff will be able in a week or so to throw these things together. We&#039;re also working on a mobile app framework and some other things like that to make it even easier. So again, the the goals of this stuff is to show so so traditionally in the government, I&#039;m sure Gail can testify to this, the way  you do a demo on data is you go out with a procurement, you hire a contractor, you get the specifications right, they play with it for a year, they build an interface. You can&#039;t have their internal formats. They don&#039;t provide those back to the government because that&#039;s now a proprietary thing. So the government gives them the data that you and I paid for with our taxes. Company spends our tax money to make it so the government can&#039;t have the data back except through their application. Right? This is this whole data gut project is about breaking through that loop. And what I&#039;m hoping you can see is we&#039;re trying to help you do that. Come to our site. Come play. come find ways to make money off it. The government would like nothing better than for some smart people to figure out how to make a lot of money off of this.&lt;br /&gt;
 &lt;br /&gt;
 We&#039;re going to be running some workshops. Evan mentioned hack and hack hackers. We haven&#039;t actually talked to them specifically yet, but I just last week got the green light. The first of the government to industry sector workshops. We&#039;ll be having four or five over the next year or two. We&#039;ll be with the media to look at how can we make metadata from these things more available so media folks who want to  find a d government data set can get their hands on it. and how will we make it so that the media can use annotation that we&#039;ll be able to get back. So, for example, if you go to one of these data sites, be nice if it said, you know, here are the stories that have been published about, you know, using this data, that kind of thing. So, I&#039;m going to stop there because I&#039;m at my time. I told him to stop me short and he didn&#039;t and let Deborah start setting up and I guess I&#039;ll take a question while she gets her laptop.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Okay. So, ties together I think the two presentations is that technically it sounds we&#039;re at the point where like you&#039;re saying web masters and tools and makers to do well.  It seems like the challenge beyond like informing the media is to inform other important stakeholders like foundation funders, like government agency funders that are funding nonprofits and other research efforts to actually make it part of their funding mechanism to create data in a way that before people before this disappears. there is an app section here where you can see some of the neat neat things that have been built only a couple of them by my lab. Most of these are by actual you know companies and things. This is the apps for America stuff. Some of that&#039;s in here but but some of these are really interesting like fly on time things. So it&#039;s not only so all the people you mentioned but we do when I say end I don&#039;t I don&#039;t expect an end user to build a mashup. I expect the end user to be able to look at one of these mashups do. And just to tell you a quick story, when data goes, so two days before the data gov relaunch on May 21st, 2010, they held a a press conference and a woman from Wired came and said, you know, I understand what you&#039;re doing and this is really cool, but I&#039;ll believe it when it passes the grandmother test. My grandmother is pretty technically literate for grandmother, but I expect to see some stuff here that you know, granny will like. Well, this made them very nervous. But about 2 days after the release, they published a paper. You can find it. Look for like grandmother and data.gov or something like that. We won, right? And the and the demo that they really that grandma liked the most was excuse me was this one.&lt;br /&gt;
 &lt;br /&gt;
 So this is the White House visitors list turned into mash data but you can&#039;t see it here because of some of the things but for example running down this side if you saw this in spar in in Firefox or Safari would be the DBPedia. So the Wikipedia information pulled through the DBPedia semantic web version about some of these people. So if you you would see a picture of Obama. So if we pick the vecundra okay so that&#039;s the vec pulled from a white house site. There&#039;s some there&#039;s some web search here news. Here&#039;s who visited him. If you click on them you&#039;ll go to their sites. You can see who they visited. we also have some social network stuff. So, grandma apparently looked at this, didn&#039;t know who VC was, and started to say, &amp;quot;Yeah, well that&#039;s okay.&amp;quot; But then it turned about three below VC was Brian Orzag and Granny was a news junkie. So the, you know,&lt;br /&gt;
 he&#039;s cabinet level, so she was excited. So she could see who visited him, notice his grandkid, I guess they&#039;re not just green, his his kids visited him. She&#039;s like, &amp;quot;This is cool.&amp;quot; So again, the end user should be able to navigate and use these apps as easily as any other apps. But but this, you know, but this app should be as easy to build as a website. And again, it&#039;s not quite there yet. But for those of you who are are sort of semly and have some background in this, pretty much everything&#039;s just offthe-shelf web technologies, PHP, things like that. I&#039;m going to I&#039;m going to stop to let let my colleagues have some time. All right. Thank you, Jim. And we&#039;ll be around all three of us will be around after to to do questions. I guess all four of us to answer questions.&lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Okay. My name is Deborah McInness and as Jim said I&#039;m another constellation chair at the room polytenic institute tetherless world constellation and I joined RPI at the beginning well at the very end of 2007 after having been at Stanford for nine years running the knowledge systems lab and prior to that I was not too far away in New Jersey at Bell Labs and AT&amp;amp;T labs for 18 years after school. Um, and so what I&#039;m going to talk about today is a couple two emerging trends that come out of Semantic Web and link data. And what I&#039;m going to do is actually first set just a little bit of context. If you&#039;ve ever seen the layer cake. So this is one of the instantiations of what Tim Berners-Lee made famous with the Semantic Web layer cake about how you start down here with kind of encoding standards and you get up to the point where you get user interfaces and applications that hopefully have trust and have some proof about what&#039;s going on. And what we&#039;re focusing on in this entire session is kind of in here. So from the the the data interchange level through encodings of some meaning and actually I&#039;ll I&#039;ll focus some on proofs and meta information about where the data came from and why you might believe it. So now we heard Jim give a nice talk and actually Gail give some nice motivation for having data on the web and if you look around at the web and its inclusion of semantics so both Jim and I have been in the semantic web since way before anybody used the phrase semantic web but it&#039;s clear that there&#039;s at least growing acceptance of semantics some notion of encoding meaning on the web And we see that trend more and more. So now that we&#039;ve got the web and some inclusion of semantics plus some notion of social spaces. So when I started as a computer scientist, I got to live in a world where I might have been the expert about the data. I built a beautiful application that was carefully constructed. It got a while to do it. I knew my users. and it was kind of a little I sat in an ivory tower of Bell Labs or university setting. But now it&#039;s the wild west. you know, the data is all over the place. There&#039;s tons of people out there interacting and so that&#039;s great and that there&#039;s more contributors to data, but now the people who are going to use that data have a little bit more of a burden trying to sift through about is it authoritative? Maybe it&#039;s authoritative, but how recent is it? So now we&#039;ve got the web and the inclusion of semantics, social spaces, and this massive tremendous potential of winking.&lt;br /&gt;
&lt;br /&gt;
We&#039;re creating a lot of new opportunities and a lot of new challenges. So, I&#039;m going to go through a couple of motivating examples that motivate some emerging trends and then my spoilers or my take-home message is that there&#039;s some next generation ontology issues that are emerging and that knowledge providence is growing in importance and I&#039;ll motivate that and then show you a path towards some solutions. So, one example that I&#039;ve been leading recently with the National Institutes of Health with their population science division is trying to help their researchers ask questions like do policies like taxation on tobacco or smoking bans restaurants barring smoking or workplace barring smoking impact health either positively or negatively and how does that impact health care costs? So if you were a researcher trying to ask those questions, what kind of data do you want to see? And and if I&#039;m the researcher, maybe I want to see one kind of data or maybe I want a different interface. And if I want to help inform the public, you know, ultimately this organization is trying to change people&#039;s behaviors. So they&#039;re basically trying to get people to stop smoking when we&#039;re working on the smoking issue and we&#039;re working on obesity as secondary things. So we&#039;re trying to get people to be at a healthy weight. how do we want to present data so that people might change their behaviors? So how do we make that data actionable and believable? And what data might we present so that people actually choose to make the right behavior changes. And then I&#039;m going to show a few pictures. And actually I&#039;m not going to do it live because I have I I rely too much on fancy graphics and IE is not adequate to show my graphics. I&#039;m gonna have to show some static things. but what kind of data do you actually want to see? And then what are the appropriate follow-up questions? So this is a very new demo where we&#039;re trying we&#039;re actually exploring right now with the National Institutes of Health about what kind of data they want to see. So one set of data that we have it has a lot of information about um oops about bands and this is the policy within workspace workplaces. The particular year that this is showing is 2007. And so you can see that some states like California would have actually expected to be better, but this is actually they don&#039;t have complete coverage of policies and work places versus some other states like New York actually does. but you can see in California when they started when the the policies started to come into place. So basically they they came in a little bit of a way in starting in 1990 and actually California led the way here although you wouldn&#039;t know that except I&#039;ve looked at the data but then in 1998 they kind of came in in a much  bigger way. And then we also want to look at whether cost actually makes a difference and whether taxation makes a difference. And we might want to look at these trends and say, &amp;quot;Wow, that&#039;s a pretty steep increase around 1995.&amp;quot; You know, maybe you want to go and look and see whether something was going on. And ultimately we one of the pieces of data that we the data sets that we have include prevalence. So how many people smoke? and actually what you want to see is smoking prevalence going down. And California is actually a pretty good state for that. And like this data point in here is kind of interesting. So you might go and say what&#039;s going on in 1995 or maybe what preceded this drop. and then what went on between 95 and 96 that it jumped back up again. So we want to help them look at the data, help them ask the right questions about what&#039;s going on. they actually put out a number of surveys to try to collect data. So we might actually also want to help them try to figure out what kind of data they should be collecting and we also might want to help them try to figure out where the holes are because one of the things that we particularly at RPI but anybody with access to the web is great at these days is the data is out there. There&#039;s a lot of data out there. So this data we see that we don&#039;t have any data in 2008 and 2009. Maybe we want to go get that. So maybe we want to identify gaps and maybe we want to identify these possibly interesting portions of graphs and maybe we want to look at questions like when I first looked at this a student put this together. I said wow taxation on cigarettes is actually going down. That seems weird. But then I read the fine print and it&#039;s adjusted for inflation. So actually it&#039;s not when you look at the data it&#039;s not actually going down. &lt;br /&gt;
&lt;br /&gt;
So, it helps you see patterns. This was the one that&#039;s much cooler and I can&#039;t show it in IE, but these motion charts. Um, so they actually when we were working with them, they said, &amp;quot;Well, let&#039;s look at a couple of states that are representative of what we at NIH consider good policies, i.e., more bans on smoking and more taxation on cigarette smoking and lower prevalence of smoking. And then let&#039;s also so some of those states were like California and a few that are off of this. And then there&#039;s a couple of states like Alabama and Arkansas that represent light policy, light taxation, and high prevalence of smoking. And then let&#039;s actually look and see what actually is happening. And I can&#039;t show this live because of the the browser problems, but but I could if it were if I had Firefox or Safari. But you can work, you can play this and then you can see it over time and you can stop it at any particular time and click over any of these circles and see how many people were smoking and what the taxation rates were and what the tax was and what the smoking prevalence was. And if you ran that and let it go to completion, you would see some states take the lead. Like California was one of the leading states if this were an action. you&#039;d see them coming out here with their ban coverage because they started everything. And then actually you see a lot of states start doing taxation on cigarettes. And you see some states like New Jersey with very significant ban coverage and very significant taxation rates. And what you can&#039;t see because it&#039;s static now. But what you really want to do is see the prevalence changing from a high prevalence, which is here indicated in blue, moving to a low prevalence, which is here indicated in yellow. And actually, you can see some of these circles change. So ultimately, if you&#039;re at the NIH, and probably many of us would like to see less people smoke. And so what we&#039;re trying to do is see is help people explore how we might get that behavior. and then we let us and them collaboratively look to see what might be causal and what might not be causal. and you might look at how things have changed over time. So the ban coverage zoomed up in 1998 in this particular state in California. and smoking prevalence, you know, went down, but it actually didn&#039;t go drastically down. But you can also see the taxation rates go up as well. So it lets them ask questions. so some of the questions are what kind of data do we actually want to display? So I started with we want to make people healthier. So we probably want to see reduction in lung cancer. We might want to see reduction in health care costs. That data is really messy. So we&#039;re not actually displaying that. But we do actually have decently clean data for prevalence of how many people are self-reporting if they smoke. because actually that&#039;s the largest data set that we have. So you know you want to go in and say well what what&#039;s the definition of prevalence and how are we measuring pre prevalence and is our data set the best data set for measuring prevalence and under what conditions do we get this data? So is this recent and how big&#039;s the sample set and are there extenduating circumstances that may or may not you know impact my use of the data and then do we need more data and do we want to make more inference. So these are just some of the drill down questions that I want to that we want to help them ask and then from a technology perspective how can we help? So Jim showed us some really nice demos and actually gave some nice perspective on the fact that we can actually build these demos pretty fast. Um, and what we really want to do is build not only build fast demos and not be not only have it so that we can build the demos, but so that this audience can build the demos and your customers can build the demos, but also make it so that those demos are understandable and so that they&#039;re actionable, so that they&#039;re potentially behavior changing. So there&#039;s some things that I as a technologist feel that I bring to bear on this. So one thing is I can help people understand what the terms mean so both in my communication of what prevalence is for example but also in trying to put data together saying if this person&#039;s using this term in this way and this person&#039;s using the same term in a different way maybe I don&#039;t want to connect it together also one enormous topic is where did that the providence so where did that information actually come from when should I rely eye on it. How recent is it? And another thing that didn&#039;t come up in Jim&#039;s talk, but comes up much more so when you look at the the longevity of these demos is how do we handle the fact that the data changes? How do we handle So NIH has given us, I think, five different dumps of the data? Do we just throw out the old data? Well, some of it was wrong and they&#039;re authoritative and that and for example, they said they had 101% coverage. Well, that seems like it&#039;s just wrong, so I should probably throw that out. But but in other cases, they said, well, this interpretation is the one that we believe now, and we had this other interpretation in the past. Maybe I don&#039;t want to throw out those previous interpretations because maybe we actually want to go back and look at that in the in the future. So, how do we handle changing demos, changing data,  different presentations for different audiences, etc. So my oops I&#039;m just looking at my time. Okay. So my two themes for my two emerging trends are one is get some encoding of the meaning and the other one is keep some encoding of where the information came from. And Marco gave mentioned that many people know me as Ms.&lt;br /&gt;
Anontology. When I went to Stanford and took over the knowledge systems lab, I said, you know, I run ontologies are us. So, we build ontologies, we maintain ontologies, we disseminate them. And what is an ontology? And I was on a panel actually in 1999 where four ontology experts came together argued for about four hours over a lot of alcohol the night before about what we were going to agree as the definition of ontology. And we essentially came up with a spectrum.&lt;br /&gt;
&lt;br /&gt;
Other people have generated ontology spectrums as well. This is one of the simpler ones that it goes from so essentially you&#039;re capturing meaning anywhere from simple level like a catalog entry so just a text string to  a much more formal very specific definition of meaning like we might encode in first order logic or a higher order logic. And right now what we&#039;re seeing on the web is relatively inexpressive. So simpleish encodings of meaning like controlled vocabulary. So usually in a single language often English to just a small amount of information saying that like maybe this this shirt is a kind of clothing. So just a simple taxonomy to shirt is made of particular kinds of materials. So I might have properties made from and I might go on to more and more descriptions more and more complex descriptions about how terms relate to each other. And when we go into the science domain that Peter&#039;s going to talk about more in the next talk, we&#039;ll see more expressive encodings of meaning. But all along the spectrum you can get tremendous value. And so actually one of the themes is that ontologies and formal encodings of meaning that are computationally understandable are gaining traction. and another theme is that capturing some encoding even if it&#039;s just a natural language about where the information came from is incredibly important. So a definition that we like to use for provenence is the origin or source from which something comes. Intention for use. Who or what generated who or what the material was generated for. The manner of manufacturer. History of subsequent owners. Sense of place and time of manufacturer production or discovery documented in detail sufficient to allow reproducibility.&lt;br /&gt;
 &lt;br /&gt;
So that&#039;s just providence alone. Then pick your favorite definition of knowledge. I picked a few just to put up here but basically some kind of belief fact or condition of being aware of something and then put them together. So where did the knowledge actually come from? And now we&#039;ve got knowledge providence. So why do you care about knowledge providence? So hopefully you got some sense of why the NIH uh&lt;br /&gt;
 researchers might care. So how recent is this data? How reliable is it? When are they going to depend on it? Um, I&#039;ve spent a career trying to make knowledge representation useful and so, uh, done a lot of sponsored research projects, interviewed tons of users, um, who essentially always say the same thing. You know, if you don&#039;t tell me why I should believe this, I&#039;m not one going to fund you, and two, I&#039;m not going to take action on the answers that your system is giving to me. So, I presume that we&#039;re going to make these viewraphs available. I just put them in there for a couple of uh articles that you  could go look at where intelligence analysts say they&#039;re not going to listen to intelligence programs unless they can understand where the data came from. Intelligent assistance users aren&#039;t going to take the recommendations. virtual observatory users that actually Peter might have a few view graphs from an joint effort that we did on virtual observatories. They want to know where the data come came from and these are all fairly well documented with user  studies showing that they&#039;re not going to accept it without that information. And there&#039;s tremendous growth in this area. So the worldwide web consortium which if you&#039;re going to look for standards on web technology that&#039;s the  first standards body I look at. There&#039;s a providence incubator group. I&#039;m on that group as well. There&#039;s some really nice documents that I really just included here so that you would have a place to go look for if you cared more about Providence. And there&#039;s a lot of quotes from all sorts of wonderful famous people like our friend Tim Bruce Lee who is famous for talking about oh yeah, you know, so if I show you something, he wants to be able to say oh yeah or why or huh. so basically that&#039;s what this whole line of research is about. So that at any time that you see something you can say why should I believe that or where did it come from. So the position that we take is that system transparency or being able to have that information about where it came from and why you might believe it and why you might disbelieve it supports understanding and trust and allows you to look at a system and and know when to take action on it. So one research goal one line of research that we&#039;re pursuing is to make interoperable infrastructure that supports this that supports explanations of everything sources assumptions answers etc. And we have built a lot of infrastructure that we&#039;ve used in a wide variety of applications from um scientific virtual observatories that if you get this picture what did that picture rely on? where was the data set coming from? if you had amazing eyesight, you would be able to see in here some metadata about the weather conditions under which that data was collected. to intelligent assistance that you can say what are you doing and why? and it says it&#039;s waiting for approval. it says what it&#039;s doing right now and what rule it&#039;s following to one of my favorites because I&#039;m a big wine and food person. to something that recommends wines with meals and why you might take that. So here it&#039;s recommending Braftoft Chardonnay for actually I don&#039;t see the question that it&#039;s being asked to recommend for but&lt;br /&gt;
 why you might take that recommendation and what information it relied on to intelligence settings of the data that it was actually coming from. Actually New York Times people might like this.&lt;br /&gt;
 So you can see the the portion of the document that it&#039;s actually relying on and then there&#039;s kind of a fancy reasoning system in the back end that&#039;s saying what it took out of that and why to taking recommendations on how to deconlict two airline paths that somebody had identified are likely to crash. So we might want to deconlict them and so this is a solution about how to resolve the conflict. So it&#039;s a wide range of applications. Oh to to prove combination. So you know might not want to just see one person or one system that believes a particular answer.&lt;br /&gt;
 &lt;br /&gt;
 But this is showing a lot of different systems that believe the same answer and it also will go and show the systems that believe the opposite of that answer. So, the thing that this all has in common is it&#039;s trying to let people look in at what&#039;s behind the scenes before they decide that they&#039;re going to buy that bankra Chardonnay or before they beat up the agent for why it&#039;s waiting, before they think it&#039;s broken so they can find out what it thinks it&#039;s doing or before you take this data set and start using it in your experiments, you might want to know more information about it. Okay. Okay. So, I just got three minutes. I&#039;m going to go through. Okay. So and these days when we&#039;re doing this from a data set coming up with a demo using data from a particular agency what&#039;s really going on is we grab that data we transform it we revise it we probably archive it we make some deductions based on it we derive these pretty demos and then what we want to do is make it so that people can see when they want to use it and when they want to combine it. I&#039;m going to skip through one of my application areas and there&#039;s backend encodings that help me also look at error conditions and oh Jim looked at this demo when in the previous talk. So, one of these applications was looking at the foreign aid and there were some questions about this demo. Um, so one of the things that you might want to do when you look at the demo is say I don&#039;t really believe that data. I&#039;m not an expert on that data, but that seems to go against the things that I would have believed.&lt;br /&gt;
&lt;br /&gt;
So, I might want to annotate it. So, hoping that an expert then will go in and look at it and see whether it&#039;s right or wrong. So now what we&#039;re doing is we&#039;re going in and annotating a number of these applications with information about where the data came from and allowing people to do things like ask question or actually this demo isn&#039;t annotated this way but one of the other ones is mashed up against in fact we should do this with our NR data against a New York Times data set that goes on in and says okay what was going on at this particular time period what&#039;s being reported in the New New York Times or any particular data set that you have access to that&#039;s going that&#039;s in this time period where I think it&#039;s interesting. So back to my original slides for the NIH where the taxation was zooming up. You know what was going on then? I finally talked to a tobacco researcher and said, &amp;quot;Well, that was the time that 46 states settled with the US government for having to support additional health care costs and all of a sudden taxation went up 45 cents a pack.&amp;quot; You know, I didn&#039;t I didn&#039;t smoke. I didn&#039;t pay attention to that, but you know, that was an interesting event that I should go back and I don&#039;t know whether we&#039;ve got data that far back from the New York Times, but we have some of the New York Times data set to see whether that would actually come out as a hypothesis for letting people look to see what&#039;s really going on there.&lt;br /&gt;
&lt;br /&gt;
So, um, a few of the points that I was making today are one, we believe knowledge providence is critical for user acceptance in many settings. Um, so we love these demos, we love this technology, but if you don&#039;t know where the data came from and how reliable it is, uh, how are you going to know when to accept it? And, um, a point that I didn&#039;t make quite as deeply here, but, we&#039;ve got a lot of data that supports having some encoding of meaning, uh, helps you interoperate with the data and helps you make the connections. And there&#039;s a reasonable amount of technology for supporting knowledge providence and ontology environments. Although it is a growing area of research and development and ultimately the open data initiative, the semantic web technology, it&#039;s changing the way I live. I think it&#039;s changing the way all of you live and I think it&#039;s going to change our futures and it&#039;s creating a lot of new opportunities. and when we put in this kind of technology, it creates even more opportunities for change and evolution.&lt;br /&gt;
 So with that, I&#039;ll take questions.&lt;br /&gt;
&lt;br /&gt;
Okay. So I have I have a question that kind of falls on the theme from the floor is a concrete example because it misleading when you&#039;re trying to draw a very simple assumption between taxation and federal.&lt;br /&gt;
 Right. Right. So this is actually a fantastic topic that we could talk for hours and hours on of and actually my work with the NIH statisticians I think is actually kind of different than the the when I work just on the rest of the data.gov of applications because I work so directly with these statistitians and they&#039;ve spent, you know, the last decade analyzing this data and they&#039;re all terrified that&lt;br /&gt;
 we&#039;re going to get that data out there and then we&#039;re going to mislead people or they&#039;re not really going to understand it. And actually fancy detailed analysis has gone into this.&lt;br /&gt;
 Plus, the data that we&#039;re working on isn&#039;t perfect. You know, it&#039;s survey data on whether you&#039;re willing to write down that you actually smoked. Um, so there&#039;s assumptions, there&#039;s weaknesses, there&#039;s incompleteness, and how do we display that in a way that&#039;s simple enough for people to understand it enough, but yet appropriate, you know, actually, uh, faithful to the methods and the data and the jury&#039;s out, you know, basically that&#039;s and but it&#039;s, um, it&#039;s imperative. So actually I&#039;ve spent a lot of time in hospitals and nursing homes and medical informatics settings lately sadly for health family health reasons. But but one of the things that came out for me in this last year of this experience was if we don&#039;t remake our so I&#039;m just looking at health at the moment. If we don&#039;t remake our education system on health and our medical informatics system and our environment, we&#039;re all in so much trouble. It&#039;s we we have no choice. We&#039;ve got to do this. Plus, I didn&#039;t say it in these in these make&lt;br /&gt;
 it in these few graphs, but the instrumented future, it&#039;s an instrumented now environment that we live in, and we&#039;re just going to get more and more instrumented. This data is coming at us at a massive rate. Um,&lt;br /&gt;
 we&#039;ve got to find ways to make sense of it in ways that might make sense to you as a PhD re researcher, might make sense to your 10-year-old who&#039;s trying to make&lt;br /&gt;
 a decision on whether to eat that donut or not. Um, and you know, kind of all levels in between.&lt;br /&gt;
 So, you know, I think these technologies can help and one of the ways that they can help is they can hide a lot of the detail in particular contexts and then they can let you drill down when you need more information. So I think this making available the context that the information was gathered in and the deductions that were done is critical. I&#039;m interested in the rights of the research and how that travels through people&#039;s use and application. Is there this is something of common how this open data can be followed and not just open data but individuals and institutions creating their own sets. Unbelievably important topic and critical to our future and and you know there&#039;s no single answer. but but we&#039;ve got to find ways to encode it, keep track of it, pass it on, make sure it&#039;s not stripped out, make sure that if you&#039;re the author, you get credit. make sure that the rights are preserved and also have access to protecting the information.&lt;br /&gt;
&lt;br /&gt;
Right. One interesting application well not just application is that it can not only be followed so if it&#039;s an effective right really good and actually I don&#039;t know whether we&#039;re gonna have time for the panel afterwards but if we do this is a great topic to follow the presentation Stephy where can we go from here and learn more about this stuff do you offer training courses in RBI where people can you know learn more in detail yes Peter Peter will talk about this. We will actually. Awesome. Okay. Thank you.&lt;br /&gt;
&lt;br /&gt;
Peter Fox:&lt;br /&gt;
&lt;br /&gt;
All right. I&#039;m Peter Fox. That&#039;s me right there in the deep south of Tasmania two years ago. Not from Tasmania, of course. And you should be worried because anytime the word revolution appears in the title of the talk, you can be sure that the fire hose is about to be turned on. And so I&#039;m gonna, you know, I&#039;m presuming there&#039;s not a strong scientific literacy in this audience, right? Anyone from a science background? Few. Okay, terrific. Um, so seriously, what is the fuss? Well, the fuss is that we have a complex Earth system. It&#039;s changing very rapidly, more rapidly than ever has been seen before. It&#039;s coupled the sources of data and information are just growing exponentially and you should all know that period bottom line and technical organizations for example Microsoft and Microsoft external research are an example of for many organizations paying amazing attention to this problem resulting in books like the thought paradigm which you can go and search on the web under did I say Microsoft yeah Microsoft research and the reason is that science is evolving this a slide due to Jim Gray from Microsoft and Alex Clay from Johns Hopkins astronomer and it basically says it&#039;s gone through four phases an empirical phase a theoretical phase and only 20 years ago did high performance computing came along and started to revolutionize how we&#039;ve done science but only in the last five years has data exploration what we normally call ecience of data science really come to the forefront and it&#039;s a synthesis of things that have come before there&#039;s advanc enhanced data management requirements. There&#039;s data mining and pushing new the need for new algorithms which has really come to the forefront. So it&#039;s it&#039;s really feeding into a lot of areas of science, mathematics, physics and computer science. But as a scientist, our working premise is this one. this is a sort of a mantra that that Deborah and I put together a few years ago. Anyone should be able to access this global distributed knowledge base and have it appear to be integrated and locally available. And if it&#039;s not, it&#039;s actually going to be very very difficult to use because scientific data, if you think that government data is complex, scientists like to do a good job of obscuring their data. But we obtain the data and information by multiple means, different protocols, different vocabularies, stated or unstated assumptions, metadata, who knows? And the red thing is what I&#039;ve been saying for about 20 years now. All data really is created in a form that facilitates its generation and not its use. That&#039;s why we&#039;re all in business. Actually, if data was created to facilitate use, we wouldn&#039;t be here. People would just use it.&lt;br /&gt;
&lt;br /&gt;
Underlying that and and you&#039;ve heard a little bit of it before is semantic heterogeneity, large scale data, complex data types, legacy systems, and so on and so on. But worse than that, there&#039;s a sort of a bigger problem and that is our two primary means of conducting science as illustrated here. The left shows the deductive reasoning approach where you start with a theory, come up with a hypothesis, compare it to some observations and see if you can confirm it. So this is sort of a typical model theoretical approach. on the right is the one that&#039;s facil not only facilitated by data because you start from the bottom with observations, you look for patterns, you come with a tentative hypothesis and try and resolve your theory. And so the way we&#039;ve built all our information systems, scientific and non-scientific, funnel people directly along those two routes. So what about abduction? And no, not the criminal meaning, but it is a methodological inference. So key, that&#039;s semantics there. Semantics introduced by Charles Sanders Peirce, which comes prior to either of these, which is to say, I&#039;m not quite sure what I&#039;m looking for. I&#039;ve sort of got an idea and it&#039;s sort of moving back and forth between the inductive and deductive paradigms.&lt;br /&gt;
&lt;br /&gt;
And so it&#039;s abductive reasoning. So again, semantics, bringing in intuition, bringing in relationships. And so this is a thing that&#039;s largely being pushed completely out of data and information systems especially for science and semantics is actually bringing it back. So to give you a couple of examples I&#039;m just going to give you two slides on sort of two paradigms of the way in which science is being conducted. One of them is in this concept of virtual observatories. So at the top here you have two real measurements the black triangle black diamond and the gray diamond. And what you really want is to know what happened where the blue diamond is. and you didn&#039;t actually take a measurement. So you virtually want to have a measurement. Now it could be done by interpolation, extrapolation from a model, from a climatology, from a whole variety of reasons, but it&#039;s virtual data. From a remote sensing point of view, you might want to take images from three different wavelengths perhaps at slightly different times and have an integrated measurement again, which tells you something more value added, the types of data integration that scientists need. But in doing this, both of these usage patterns exacerbate our data management problems, the challenges, and now we&#039;re managing virtual data sets as well. So this is one of the big problems and and this is the revolution that&#039;s going on. We didn&#039;t know how to manage regular data sets. Now we&#039;ve got virtual data sets. So about eight years ago, Deborah and I started on a project to unlock these virtual observatories. This is a little animation of a a virtual observatory slide to show you how we&#039;ve implemented semantics in a production system. It&#039;s been in production since 2007. It&#039;s had evaluations. It&#039;s won best papers, all sorts of things. So here are all the different data sources. And up the top are things like virtual observatory portals, web services, and APIs. And in the middle is this semantic mediation layer which includes the ontologies capturing all the important scientific concepts, parameters, instruments but also descriptions of the data and the service classes as well. So there&#039;s a full ontology behind this. It maps the queries to the underlying data mediates across all these repositories. But the most important thing really is it introduces this interaction, this smart interaction so that you can start offering up alternate hypothesis, make suggestions in an openw world setting rather than assume you know everything and funneling people down the inductive deductive route. You push things up in front of them. so it allows this extra interaction. It gives semantic interoperability well- definfined meaning between each of these modes of accessing this data and it allows you through higher levels of mediation higher levels of of ontologies mapping to non-speist vocabularies educational vocabularies and inventory vocabularies and you can and you can go right up the stack and you know the business case here is where does semantics add value? Well, it adds value in unlocking these data resources. It adds value in inter operating with web infrastructure between them and and among between the cataloges and it gives you the ability to push it into other domains. The second piece which also is known as a VO is a virtual organization and this is the way a lot of scientific collaborations are being conducted these days. A virtual organization and you probably participate in these. It&#039;s a group of individuals who may be&lt;br /&gt;
 geographically dispersed. They may come from different institutions but they use the internet largely and tools to come together under a common goal. That that co goal might be short-term but it&#039;s often the long-term common interest and the is information technology and most importantly if if you&#039;re in a virtual organization this is your primary takeaway. When you work in this context the role and responsibility and status you have in that virtual organization may be completely different from your home institution. And that in fact not recognizing that can cause significant problems. And this is important for scientists as they start to deal with these larger amounts of data in a completely different mode. And so a virtual organization has coupled elements of technology some organizational structure and communication. So that to facilitate massive collaboration which is a subtitle of my talk, you have to take this into account. And because they have a high degree of informal communication, lots of little emails, you have to basically instrument the infrastructure by which they communicate in in the systems that you design and attach them to the data in in exactly the same way that we we&#039;ve been hearing about for Providence. yeah, so a method that we that we have developed over again about now six years. So it&#039;s a modern informatics approach. It takes advantage of the scale-free nature of the internet and semantic networks.&lt;br /&gt;
 &lt;br /&gt;
Scale-free meaning that the infrastructure that you build will work on small groups for small problems and large groups from large problems and in between it&#039;s a log log distribution and it involves fundamentally use cases all the stakeholders in the virtual organizations distributed authority full recognition of access control ontologies to mediate things and maintenance of identity. And this particular circular and iterative development cycle that we have which is modeled largely on the software development cycle starts with the use case proceeds around export review and and most importantly it only adopts technology late. It does several iterations before it goes and even looks at technology. So it&#039;s not technology bound. It&#039;s meaning bound. And we rapid ro prototype open world iterate redesign redeploy and this we teach to graduate students 13week course. Deborah and I developed the the course this this year Joanne and Deborah are teaching it and it&#039;s called semantic e science.&lt;br /&gt;
 &lt;br /&gt;
So let I just want to give you an example of a use case and they&#039;re not trivial use cases. So one of them is in marine habitat change. So this is an image that&#039;s taken by a thing called a habitat camera courtesy of Scott Gallagher at the Woods Hall Oceanographic Institution. This is interesting for many reasons. These cameras are towed behind commercial fishing vehicles off the Atlantic seabboard. They obtain about a terabyte per run and they&#039;re analyzed for scolop shells, size, shape, color, place, number, and density, fragments. But other people want to use this data as well. So what&#039;s this?&lt;br /&gt;
 Sometimes the dirt and mud, what&#039;s considered noise to the skull counters is another person&#039;s signal. They want to understand the sediments and the rocks that are there. And the use case might be what&#039;s the temperature and salenity of the water and is this marine specimen is that meant to be here? Is it a flora? Is it a fauna? And is it part of an ecosystem change? These are the things that we&#039;re starting to answer today&lt;br /&gt;
 and facilitating things like integrated ecosystem assessments, real applications. So, it&#039;s a similar diagram to the one you saw before. Rich scientific data repositories, software applications and tools and integrated applications used by a variety of stakeholders. The Nature Conservancy, Noah Marine Fisheries, and the US government&#039;s council on environmental quality. And notice in the gray boxes,&lt;br /&gt;
 it&#039;s semantic web in action.&lt;br /&gt;
 &lt;br /&gt;
Largely vocabularies leveraging things like international standards, organization standards that are accepted worldwide. And so we&#039;re starting to to to build and deploy these types of systems. And what you can then get is a generalization of this where our semantic web implementations are moving away from simple data integration to application integration frameworks. So we are able now to mediate these types of application level mashups that you&#039;ve heard Jim Jim talk about and ultimately these are used by people a broad variety of people. So done yet but there&#039;s a problem and the problem is that this is what we would normally consider the full life cycle of data.&lt;br /&gt;
&lt;br /&gt;
It&#039;s the thing that&#039;s sort of held up as you know we go from data we convert it into information which people can look at and then it becomes knowledge we might write something about it and then we really learn something about it now I&#039;m not going to go through this slide when I teach this I teach a course called data science I spend 30 minutes on this one slide going through it in excruciating detail but the problem is this is not actually the full life cycle of data it&#039;s a micro life cycle of data because anytime we add provenence anytime we add supporting information Any time we add context, we&#039;re adding knowledge. And so you know what real data pipelines look like.&lt;br /&gt;
&lt;br /&gt;
Ready? They look like this. And this is the one without the animation. And so raw data is up at the top. Highly integrated data products are at the bottom. On the right are all the people and feedback loops that might occur. On the left are all the people and processes and metadata that might get added. This is the reality. This is the reality of the world. And all the data that you&#039;re seeing on the internet comes any from anywhere along this line. But the problem is, and you can see the text on the left, there&#039;s fragmentation, there&#039;s disconnection, there&#039;s encapsulation, and all of them are bad for what I call the illusion of transparency that we&#039;re trying to get to in open data. It&#039;s an illusion. And you can come and ask me afterwards what word I use to replace it. But the ecosystem of what scientists want. So transparency is I forgot to put it in quotes here. is part of it but its accountability, its identity and what is it is made up of are these elements in the middle and some of these are in the semantic web stack and in fact this is all these are all provenence elements and that&#039;s why it&#039;s so critical debased we all we all are emphasizing it saw this definition which is the one that we use but provenence is only part of it because in sort of answering these scientific questions determining fitness of purpose all these things the knowledge base that we use is is multicomponent provenence is one domain science of any particular science is another and because there&#039;s data processing you have to have that in there described as well and so we construct these knowledge bases and I&#039;ll just give you just a little example of two of them so this is a providence aware faceted search that we&#039;ve developed in the area of solar physics where the facet boxes are along the top and maybe you can read them. Um, people can select they can put these facets in any order in any combination and the the relationships between these facets determine whether you get results from the queries or not. Um, and when you get a and there&#039;s a provenence in in here because it says the cloud cover is clear even data products. Up comes the image and here&#039;s the full providence trace with the ability to sample the metadata, find out that the cloud cover is in fact clear and you can drill down through all of this. Scientists actually really like this. Now you&#039;re starting to get this explanation, justification, verification. We&#039;re not proofing, proving, and trusting anything in this particular example. One more. Oh yes. So this is data, provenence, ontologies, RDF, RDFA, and sparkle. Another one from NASA. And this one&#039;s actually fairly cool. So, it&#039;s not going to be an atmospheric science lesson, but NASA has lots of satellites, has lots of instruments, and what people like to do is do correlations between those different instruments and different satellites to do sanity checks. So, just look at this one. This these are two different instruments measuring the same thing on on two different instruments on the same satellite. Take a look at that.&lt;br /&gt;
 &lt;br /&gt;
Does that look right? It&#039;s a correlation. So, good correlation would be red, orange. Negative correlation is blue. Take a look at this one. So, these are same satellite same instrument on two different satellites. Look at this. What&#039;s that? You probably can&#039;t see, but you know what this is? That&#039;s the international date line. Does the atmosphere know about the international date line? Do satellites know about the international date line? No. So what&#039;s wrong? Welyahoo&lt;br /&gt;
yahyl, you can explain it to them at the end, but what we&#039;re doing now is on the fly, we&#039;re intercepting these selections and because of the provenence. So what is that? What is that? And we have a thing called a semantic advisor. So that when they make these selections of parameters they want to look at for a particular time and how they want to visualize it.  So here&#039;s the the thing that they visualized. It brings back a chart that says this is the first one you want to look at. This is the second one you want to look at. And in this column is are they different? Well, it says okay the data sets are different. The platforms are different. The time they cross the equator is different and one is one is ascending and the other is descending. And your definition of a day is the same. But when combined w with these other factors, you produce an anomaly. and we can actually explain it to them and correct it for them. Okay, that&#039;s pretty cool. What&#039;s the anomaly? What ah the anomaly is the difference that&#039;s introduced by the same day-to-day definition, but these two factors here, the fact that the equatorial crossing times are different. So, they use the same definition of the day, but the actual equatorial crossing times are 5 hours apart. And that&#039;s produced and this is an overpass time difference which shows you exactly why the anomaly is occurring and the inas that&#039;s just a just a blow up of it so you can see and this is what it looks inside it&#039;s distributed so NASA god generates the user request I won&#039;t go through this in in any detail but it uses owl it uses domain ontologies it use general rules it uses inference it generates an advisory and it sends it back to the other And I can show you the general rule.&lt;br /&gt;
 &lt;br /&gt;
It&#039;s about 18 lines long. So these are the things these are the types of things that we&#039;re building for real agencies um with real problems using very modern ontology development techniques and and tools.&lt;br /&gt;
 All right, almost done. Now some animation. So here&#039;s this diagram. I&#039;m using it for a different purpose here. So at RPI, since the three of us have been there, we&#039;re in fact developing curriculum at the graduate and pushing some of it down to the undergraduate level. And so in the yellow ellipse is the area of data science which I teach. In the middle orange is what I call Xinfirmatics. You can ask me about that later. which I also teach and Joanne and Deborah teach the semantic e science which cover this this space here with these overlaps and I didn&#039;t put on here but underlying it all is web science which Jim teaches so we&#039;re trying to train people on so on day one when they graduate they can come and do all this stuff for you or someone you know guess you can&#039;t teach wisdom no don&#039;t want to try it just yet need something else to do in five years when we solve all these problems.&lt;br /&gt;
&lt;br /&gt;
That&#039;s a joke. So here&#039;s my summary. semantic ecience approaches are changing many fields of science extremely rapidly. This informatics approach enables integration at a variety of levels. So virtual observatories bring the data together. Virtual organizations bring the people together. And we&#039;re finding massive, you know, massive can be a relative term here. collaboration through vocabulary  mediation because that&#039;s how people collaborate. That&#039;s how you get that&#039;s how it works. But it exposes many issues. Transparency and the dependence on provenence and the new things that are popping up like quality, fitness of purpose and trust and how you make those computational. And so what we&#039;re doing in from my point of view is is exploiting this tension between production and applications and  research issues and implementing real things. And my goal personally is to restore abductive reasoning to the conduct of scientists. whether you&#039;re a specialist or a non-speist and that is going to be done with semantic web in an open world environment using the internet as your primary computer and I&#039;ll just finish with that slide that shows the the constellation thanks 20 minutes do you want to take some questions or do you want to yeah So for for solving really complex&lt;br /&gt;
 problems, how do you choose what data to use as a source?&lt;br /&gt;
&lt;br /&gt;
Good question. So very very quickly here I&#039;ll just pull up this diagram. So this team of people here, this small team includes the scientists, it includes people who are familiar with the data as well as well as software engineers, knowledge representation people. object modelers and so on. So they&#039;re in it from the start. So they identify so there&#039;s two modes. You identify the obvious data sets which you saw at the bottom of some of these and then increasingly we&#039;re going out and finding data sets because data sets are now being broadcast using things like GSS and Atom. And so we&#039;re going out and finding those as well and if they have the right markup we can discover them and see that they&#039;re relevant. Peter, can you come to your last?&lt;br /&gt;
Y. So Marco just asked me to take one minute to sort of answer the question he asked before sort of what next, what do we do, etc. So to start with, of  course, we&#039;re a university group. For those of you who don&#039;t know, RPI is in New York State. We&#039;re not in New York City. We&#039;re about two and a half hours up the road in Albany, Troy area. Not in Rochester. Not in Rochester. That&#039;s a lot of people think we&#039;re Rochester or Rochester Institute Tech. We&#039;re RPI. Renelers. we offer the traditional stuff that academic organizations do. So we have undergraduate graduate programs. Anyone looking for a master&#039;s degree, we have an information technology and web science program which has a lot of this stuff in it. We&#039;re in the process of creating a PhD program specifically on this stuff. We just can&#039;t figure out what to call it yet. We&#039;re we&#039;d love your ideas if anyone who would send anything. We also, as you probably have noticed, are not like a traditional academic thing. You didn&#039;t&lt;br /&gt;
 see slides full of mathematics and things. We&#039;re very interested in the whole transition of technologies, building technologies.&lt;br /&gt;
 obviously have a lot of semantic web background flavor. I actually am getting bored of that. I&#039;ve been doing it for 20&lt;br /&gt;
 years. So, so I want to do web 5, but I used to say 40, but now you&#039;re pushing me. So, but anyway, but so there&#039;s that. But again, you know, we have a lot of opportunities. Um, we don&#039;t do any real training courses. We&#039;ve been helping people develop training courses. We&#039;re very interested in working with people who who&#039;d be interested in taking this stuff further. In a sense, what we&#039;d&lt;br /&gt;
 really like is is to stop doing the stuff that&#039;s evangelizing this technology and go back to what we were doing a decade ago, which was being the lead. And I think Peter&#039;s shown you and Deb sort of what some of those leading edge things are. So in a sense we&#039;ve sort of you know played with this semantic web stuff a long time are moving in some of these new directions very interested in finding partnerships&lt;br /&gt;
 across that whole spectrum. So you know you have our contact information on the mashup stuff. We love talking to people and stuff like that. Marco wanted a&lt;br /&gt;
 second after and then I think he was going to throw it to questions just at this point interactive conversation parties meeting around&lt;br /&gt;
 Definitely would like to know just thinking I mean why don&#039;t you mash up everything that we spoke about perhaps have you considered maybe in data dog now that we have this all these curated data sets having some computationally you know way of doing like cherry picking through okay you know what I&#039;m I live in New York City I log into data go I put in the related data sets and you know what this might be something something I might know that somebody in Washington DC doesn&#039;t know. I can then add some more. Right. So I actually didn&#039;t get into that whole thing. that&#039;s a big part of what I&#039;m actually moving into this. So and and the science stuff is a particularly good domain in this. So policy makers make decisions based on this data. Those policies affect our lives. Yet we don&#039;t really have any input back into the data. And the qu and and one of the things that&#039;s really fascinating when you talk to the CIO and CTO of the US and people like that, they want to turn it into a conversation. And the question becomes, what does it mean to be a conversation? So they&#039;ve got a site you can get to from from data gov that will let you suggest data sets, make comments, tell them about your apps. Anyone who develops an app off this stuff, you can, you know, ask them to put it on their page. But the data itself is is something you know it&#039;s a lot of stuff. How do we start talking about it? How do we come? How would how can you say what&#039;s going on here? So one of the things that&#039;s really interesting in these visualizations is we&#039;ve been going back to to government organizations and saying hey guys you&#039;ve got a bug in your data and they&#039;re like how do you know? And I&#039;m like look right you we got I didn&#039;t show all these examples but like we have wildfire data you know how many acres were burned? How many fires in 1985? The answer is zero. and we&#039;re pretty sure there were wildfires at IA5 and that that&#039;s a data error. So we showed that to the you know so so the guy who now runs data gov is the former CIO of the department of interior who&#039;s moved over to data gov and he says one of the reasons stuff like that started convincing him that this is good for the government. This is how they will know more about their data. one of the reasons we we did this mass was so that they could do some of these visualizations before they release the data so they can find some of their own errors. I have I have a whole bunch of anecdotes I won&#039;t go into. When you start getting into the scientific data much, you can&#039;t do that by eye, right? You can&#039;t look at, you know, some of the stuff Peter&#039;s talking about and say, well, you know, of these million observations, we think these 37,000 are somehow anomalous, right? So, so you really need to start looking at what are the tools, techniques, etc. Uh, and that&#039;s also why you need this whole infrastructure. Where did the stuff come from? So, so again, yeah, you&#039;re you&#039;re mashing up all our stuff and in fact that&#039;s part of what we came together to quick follow on to that the So, if you&#039;re going to get massive collaboration, you have to solve the last mile problem which is not to computers from computer to person and localization. Okay, it&#039;s familiar things that they can actually do that are actionable fundamentally important and you know the web is making that much easier. And how do you make it make a connection to them? Um, so we we might know something about who you are. So you said your location. So that&#039;s one thing about who you are. But when we might have additional context and if we can tie into something that you understand, you&#039;re more likely to understand the rest of our message.&lt;br /&gt;
 &lt;br /&gt;
 Yeah. I&#039;d like to go back to the question of intellectual property this open data and so on. I&#039;d like to ask you know what kind of is it truly open public domain data or what kind of pitfalls should we look for&lt;br /&gt;
 example NIH has the UMLS you know 100 vocabularies but you know a number of them have their own licenses and so on so what happens with this data is truly public domain can be used in any country and so forth so so data gov any data set they release has to be under what&#039;s called the data gov release policy which basically says can&#039;t reveal any personal information can&#039;t reveal anything that you know against the US national interest and other than that anybody can use it for anything so it it&#039;s very open now in the Brit British case they&#039;re actually releasing some data sets that have some of those restrictions so they&#039;ve spent a lot more time looking at issues like the creative so so it&#039;s easy when the data lives in a document ument to put a tag on that document when the data comes out of the document how do you put a tag on the data right you know how do you how do you say this if I mash up this data with that data how do I know that&#039;s where a lot of the prominent stuff comes in so there&#039;s still a lot of hard problems down the road the US has found one solution which is simply release the stuff that can be released the UK has been using some different things some of the other governments so pretty much what the governments are doing is they&#039;re giving away stuff they&#039;re pretty sure is okay no matter what you do with it. And that&#039;s limited some governments to give away much less stuff than if they do it. So we think as technology solutions help put policy restrictions on things, explain things, bring things together, let you know where things come from, that may help. But right now, at least for data gov, everything there, all 273,000 data sets are available for you to do whatever you want with in whatever way you want. And&lt;br /&gt;
 in fact, they encourage you to Gail said towards the end something about FOYA requests, freedom of information act. So there was a a White House blog about three months about two months ago where they said since they&#039;ve created data gov, the amount of money spent on foyer requests by the administration has gone down by hundreds of thousands of dollars because people had to ask for this data. They&#039;re like if we have to give it away, if they ask for it, why don&#039;t we give it away first? Now what&#039;s the problem? Search.&lt;br /&gt;
 &lt;br /&gt;
 Right? How do you find the data you&#039;re looking for? Right? So, so it&#039;s, you know, the a lot of the FOYA guys are saying, &amp;quot;Can we have this information?&amp;quot;&lt;br /&gt;
 They say, &amp;quot;It&#039;s in data.gov.&amp;quot; And they said, &amp;quot;Where?&amp;quot; Government&#039;s like, &amp;quot;Oh.&amp;quot; So, but that&#039;s not true in in all areas.&lt;br /&gt;
 So, it&#039;s true in the data.gov, but in some areas, uh, you do need to put some policy information in there and then we need to check for compliance. So, if I&#039;m, um, allowed to give it to you, uh,&lt;br /&gt;
 if you&#039;re required and I wasn&#039;t procluded from giving it to you, but then if you&#039;re going to be required to give it to somebody else if they ask for&lt;br /&gt;
 it, then I might need to limit my ability to give it to you. So, we need to encode the policies and then check the compliance separate research uh&lt;br /&gt;
 efforts. We&#039;ve got work on encoding those policies and checking to see whether they&#039;re complied with question about linked data. So, two&lt;br /&gt;
 different two same question for two sides here. can you give me some examples of linking data sets on data gov and then also on the e science side&lt;br /&gt;
 again linking data within the data set not a mashup outside the data set within data sets themselves right. So we&#039;re doing several things&lt;br /&gt;
 in that space. So first of all there is when you do the mashup against stuff that&#039;s common terminology. So for example one of the Google visualizers&lt;br /&gt;
 is a timeline. So I got several different data sets with time information in them. I can pull those together even though they weren&#039;t intended. So the simplest kind of&lt;br /&gt;
 linking is common naming over stuff we know. So one of the things the governments are very interested in and there&#039;s actually an international effort&lt;br /&gt;
 starting up in this is you know what are some some specific kinds of terminologies for those things. Second of all we do a lot of linking to the&lt;br /&gt;
 DVPedia to to census data things like that. So we&#039;re actually all our data once it&#039;s an RDF now has URIs. Those&lt;br /&gt;
 URIs could be linked to other URIs by either the complex mechanisms owl same as by procedural code by I think so in&lt;br /&gt;
 in my world the sort of broad world we&#039;re trying to solve it without deep reasoning we&#039;re trying to sort of say somebody sticks that information in&lt;br /&gt;
 there and everybody takes advantage. So that&#039;s sort of the linked open data cloud and our six billion triples are tagged and the 13 billion is not enough&lt;br /&gt;
 in-n-out links and it&#039;s true because most of our links are to the data set level not to the data element level but&lt;br /&gt;
 of course the data element is 14 if you don&#039;t know what it stands for you can&#039;t do much so we&#039;re we&#039;re very involved in thinking that stuff through but again my&lt;br /&gt;
 world is a very broad very shallow world Peter the other well I would say on on the scientific&lt;br /&gt;
 side there&#039;s not an enormous amount of what&#039;s&lt;br /&gt;
 termed as linked data yet the agencies are starting to to work on it I would consider the work that we&#039;ve done as&lt;br /&gt;
 very linked data and so we do data integration within data sets all the time you can ask questions like what&#039;s the state of the atmosphere and it knows&lt;br /&gt;
 to go and get temperature pressure and density and composition without you even having to ask it so there lots of smarts and these are ontology using you know domain expertise.&lt;br /&gt;
 So these examples I think are more complex understanding and then integration. so I think actually&lt;br /&gt;
 that&#039;s kind of one end. and then actually in the NIH end we&#039;ve kind of kind of in the middle where we do a little bit less understanding and a&lt;br /&gt;
 little bit more integration. So we&#039;ve got ban by region, ban by state, ban by county, time frames that don&#039;t&lt;br /&gt;
 quite map perfectly and then we have to basically do a translation step so that they can actually be&lt;br /&gt;
 integrated better and and you know so the fast answers there&#039;s a continuum and what you can see was fun about bringing our lab together is we want to figure out how we stop&lt;br /&gt;
 these being different communities that don&#039;t talk to each other and really start to understand across the whole thing. But that&#039;s our scientist.&lt;br /&gt;
 &lt;br /&gt;
Yeah. Yeah.&lt;br /&gt;
Yeah. I was wondering one of the most contentious issues is the so-called social graph essentially a lot of personal ontology and so this to me is a big problem that companies are monopolizing people&#039;s identities and their networks and that will place severe limitations on your ability to integrate all your personal data. is there any open way to to solve this issue and you know establish the unique identifier for me that I can link&lt;br /&gt;
 to all the other data right well so so trick number one is establish a unique identifier for you when you want but not that&#039;s always unique right so if you go to this system and this system and you don&#039;t want it known that you&#039;re the same guy you also need some so so what I&#039;ll tell you is there&#039;s a lot of thought going into that we work with I&#039;m I&#039;m working now with the guys who are doing what what&#039;s called web ID. So these were the the guy who who sort of has pushed the be the idea that&#039;s becoming best known was the chief architect at Twitter where they had exactly this question all over the place. Right? I got this Twitter persona. Twitter built itself as an open thing that wanted to link stuff but it can&#039;t because they can&#039;t touch into this and into that. So, so web ID, um, the idea is to have something sort of similar to an email address, but reverse&lt;br /&gt;
 indexable to something on the web that when you look at it would say, here&#039;s what I&#039;m called in Facebook, here&#039;s what I&#039;m called in Twitter, here&#039;s what I&#039;m called in email, that kind of thing. So,&lt;br /&gt;
 in in in sense doing that and then the next step beyond that, something Berners Lee and I have been playing with a little bit is groups. So almost every&lt;br /&gt;
 application you use, you have some kind of group model. In email, it&#039;s a mailing list. Facebook, it&#039;s a Facebook group.&lt;br /&gt;
 &lt;br /&gt;
Twitter, it may be friend groups or or there&#039;s a lot of other, you know, kind of informal mechanisms on and on and on. And so the problem is I&#039;m I&#039;m here now. And someone says something or shows a URI and I say, gee, I&#039;d like to share this with everyone in my lab, right?&lt;br /&gt;
 Well, now it would be application dependent. Who in my lab is in that application? How do I know all their addresses? We want to break that stuff2:00:012 hours, 1 secondopen. So what you&#039;d like is sort of an O. So the original idea was open social and that&#039;s sort of moving. But the newer idea is open ID like stuff but with real meaning behind it that lets you control that linking and also lets you deny that linking. Right? So if you don&#039;t want me to know where you&#039;re who you are in Facebook, you should have an easy way to just not have it show up in that stuff. And that&#039;s a little bit problematic because Google wants everything in there and Yahoo wants everything in there that Google doesn&#039;t have. And so, so again, there&#039;s a lot of IP business things like that. But I&#039;d say that if you&#039;ve been following this wired the web is dead, long live the internet. No, the web is wonderful. which I actually think is a stupid debate, but I&#039;ll I&#039;ll take that offline later. But a big part of that is are we going to move to an apt kind of world where everything is separated? So&lt;br /&gt;
 we&#039;ll silo everything or are we going to find a way to put it together? And where&#039;s the place to put it together?&lt;br /&gt;
&lt;br /&gt;
It&#039;s in that semantic backbone of the of the web. And that&#039;s really where a lot of the I think the the the cool stuff happening below the infra, you know,  really deep in the infrastructure. The guys who are doing this are are, you know, deep webbies from the old days who are really looking at, you know, what what&#039;s the right kind of naming conventions, how you do this. So, it&#039;s pretty cool stuff, but I think we&#039;re still a few years from seeing it really break open. Yeah. kind of in following up on what you just said, I see restful style, I see restful style solutions as an important way to make this stuff happen and I didn&#039;t see it in your instead. So, you know, I mean, is it just starting to make sense in this context or is it, you know, I&#039;m going to give you a fast answer and and be happy to talk for hours after it if you guys have other opinions, but I&#039;m sure, but what I&#039;d say real quick is that the web has become a development platform.&lt;br /&gt;
 &lt;br /&gt;
 In fact, one of the things I think slowed the semantic web hitting as big as it as we thought it would, as fast as it would. Now it&#039;s finally turning some of the corners we after 10 years that we predicted after five. And a lot of it was the realization that we had to get back into that infrastructure, right? Really change when you change the labeling of web link, something big is happening, right? You change the meaning from a tree to a graph. Big stuff. So, so again, a lot of stuff had to percolate up and there&#039;s actually some pretty cool stuff down there now. So, what I&#039;d say is a lot of this stuff really does fit with REST. a lot of the new sparkle protocols, a lot of the sparkle 2 stuff is really one one I guess they call it now is looking at that direction. So, so you know, we&#039;re firm believers in it, but there&#039;s also a group that says, well, you know, now you&#039;re also seeing kind of this third. So, you&#039;ve got the internet level, the web level, and now you&#039;re starting to get this new API which says abstract up the information levels, right? So, this the web is all about making it so you don&#039;t need to understand how the internet works when you&#039;re developing an app. So why should you have to understand how the web works when what you&#039;re trying to do is say get the data that Deb&#039;s using and the data that Peter&#039;s using and show it to me on the timeline that that Google gives me right why do I need to know anything about RDF and that should all be sort of somehow you know back there in the infrastructure there&#039;s a lot of discussion of the questions you&#039;re asking I don&#039;t think there&#039;s a simple answer but if I&#039;m a provider we sure hope you do it as either REST or a simple service API or an API So, so, so you know the example I&#039;d give you is is visualization has for years been the long pole in the tent of doing anything to data, right? I got this data, I process my data, I run it through some data analysis thing and now I can publish a journal paper that says I ran my data through the data analysis and the number is 63. And then if you want to do any kind of demo of that and stuff, you had to hire a whole new team of programmers and a whole new we&#039;re doing it all through APIs now and XML transforms and things like that. So, so again, this web level of of abstraction is doing some very cool stuff in that space and and that&#039;s pulling some of the rest of some of the services. I I won&#039;t go on all night. I could I was about to say I could go on all night and Peter would say you are about corporations who might have had very comfortable arrangements in the past with government data. Do they fight or is that an issue at all? Do you want to talk?&lt;br /&gt;
 &lt;br /&gt;
 I feel like someone else that question has that been in the news recently about an oil spill. It&#039;s it&#039;s there. And so the short answer is when these types of revolutions in how we do things come, people don&#039;t go quietly and it it takes a while for these these changes to propagate and we&#039;re seeing resistance. But and that&#039;s why a lot of there&#039;s a lot of initiatives about openness and claims of transparency because the more you make people aware of it, people started to get interested in it and start to hold those those entities whether they be companies or even governments accountable and accountability was on my on that slide&lt;br /&gt;
 there. And that really is starting to change at the international level and this is true for science and medicine.&lt;br /&gt;
 there&#039;s a worldwide effort to to push these open data open data policies everywhere. Absolutely everywhere.&lt;br /&gt;
 &lt;br /&gt;
 And as their resistance like you would you know it turns out by luck that one of the best things about the semantic web was we weren&#039;t as successful as we thought we&#039;d be as fast as we thought we&#039;d be and I can give you a long discussion because in a sense what happened is a lot of people were attacking other things where they thought there were more money while we were getting our act together. And now all of a sudden, you know, I&#039;m reading papers that say, well, the semantic web will never happen. And I&#039;m like, you ever hit a like button, you know, that&#039;s RDFA. I mean, that&#039;s stuff that came right out of those early papers said exactly how to do it. I mean, so, you know, this stuff is really now coming out. And I think because it&#039;s coming to the infrastructure level, it&#039;s hard to keep it out. And that was a lot of how the web worked. So, we cross our fingers that what I just said will keep going.&lt;br /&gt;
&lt;br /&gt;
But well, I think I think like the and actually something that was on the cover of the times recently healthcare again with Alzheimer&#039;s and the searchers and collaboration around that effort are all positive things. So, but I&#039;d like to ask a slightly more technical question about how you guys are moving forward with interoperability typically in science and with graph databases and the kind of graph programming that are at Sandina is pretty so we&#039;re starting to do a lot with graph graph based importance and even a lot of the algorithms that are coming out now are graph based algorithm because in in science especially I didn&#039;t go into in great detail. you know a lot of the algorithms especially on the semantic web scale are n^ squ n cubed.&lt;br /&gt;
&lt;br /&gt;
Okay that&#039;s terrible in mathematical terms. You want to be n login and even we&#039;re talking now about sublinear and that&#039;s the only way we&#039;ll be able to keep up and currently all the techniques are based on graphs partitioning you know trimming the graph all these all these types of things. And so the intention really as soon as was talking about APIs, the way you deal with it now is that such a low level of programming algorithm that has to be raised up a very very significant level. Otherwise, no one&#039;s going to be able to create apps to actually use these, you know, find results from these data or use these applications. And so we&#039;ve got a whole API evolution that needs to come along.&lt;br /&gt;
&lt;br /&gt;
I&#039;ll leave how many more we take. Yeah, we have another five 10 minutes. Okay. I would go with a question that&#039;s a little bit opposite. You talk about open data, but a lot of data is not open and actually there&#039;s some reasons good reasons for that data not to be open and be hidden in websites because you don&#039;t necessarily want to reveal that. Now, what is the is there work that&#039;s being done behind that that goes into the security perhaps a security layer for Oh gosh. Yes. Next question. So, so first of all, there&#039;s a lot of enterprise stuff going on. Obviously, it has to think about that. Second of all, there&#039;s a lot of work I mean from day one. Remember, a lot of this stuff was funded out of DARPA right at the beginning and was about government interrupt. So, you know, it&#039;s not coincidental data gov uses semantic web stuff in both the US and UK, the biggest projects because designed to do that stuff and it&#039;s just we thought it would be within the government they did it and we had to get them to give it out first. But but a lot of that SEC so so there&#039;s a lot of security level there&#039;s also a lot of stuff that&#039;s so what was really nifty is I was at you know something in a community that&#039;s not doesn&#039;t like to talk and you know I have to sign papers I&#039;m not allowed to tell anybody and bet anything and you know they&#039;re asking exactly question I said well you know you already have a web security model right and they said yes I said okay you&#039;re done right I mean to one level of approximation this is just web stuff now when you start talking about data aggregation, data propagation. I mean there&#039;s a lot more there but from just simply the locking the door doing the access control the whole web is hard to do that for but the semantic web doesn&#039;t make that much harder right it does add some new and interesting issues and so there&#039;s a lot of us working on it that I mentioned we were doing this work on policy stuff that we&#039;ve been doing that for six or seven years now with the MIT group that Tim runs and Tim Ernestly runs and it&#039;s all about policy awareness how do you make policies is explicit but but how do you make things accountable so so the current security framework is if I can&#039;t protect it completely don&#039;t put it anywhere of course the new zeitgeist of the digital native is we want to share everything so the question becomes how do you build a&lt;br /&gt;
 security and privacy and control layer that&#039;s that&#039;s usable by people who want to share not by people who want to hunt&lt;br /&gt;
 and so there&#039;s a lot of interesting discussion. There&#039;s been a lot of workshops and and I won&#039;t say there&#039;s anything, you know, I can point it at and say, &amp;quot;Yeah, there&#039;s a solution right&lt;br /&gt;
 there.&amp;quot; But I&#039;ll say that that&#039;s very active and exciting topic.&lt;br /&gt;
&lt;br /&gt;
And then there&#039;s also work that looks at if you think you want to share everything, should I inform you about the potential consequences of sharing everything and should I provide some infrastructure that allows you to protect things that maybe you should have protected, but you didn&#039;t think about it. So there&#039;s also kind of this this ethical moral kind of slant to all of this that not only are we&lt;br /&gt;
 providing technology providing and designing technology but we&#039;re also trying to think about the ecosystem that that should sit within and being&lt;br /&gt;
 responsible creators knowing that some of the people who adopt it are just not going to have thought about all these issues that they maybe should. Well, I&#039;m&lt;br /&gt;
 thinking for say for situations like say healthcare, right? I mean, right now people get up in arms about well Facebook is giving away too much of my information. But guess what? Your&lt;br /&gt;
 healthcare information is going everywhere and what sort of control are you ever going to have on that? And if you do, will that would there even be&lt;br /&gt;
 infrastructure for well I might be giving a control or permission to this company, but now this company has all&lt;br /&gt;
 these relationships with all the other companies or all the other sources that might be using that and how how does that marking or that&lt;br /&gt;
 data traverse to the entire semantic web and however it&#039;s being right right to so policy is part so a&lt;br /&gt;
 transparent policy is a policy that can be encoded and enforced is part of it um but another part of it is if I if I give you all the knobs and bells and whistles to protect it can you do it and can you even anticipate what you should be protecting so healthare care information is kind of the first example. So, HIPPA is probably doing you a favor, but sometimes it&#039;s not. You know, sometimes it&#039;s keeping information from your doctor or your healthcare provider in a timely manner that could have saved your life. So, you know, then it&#039;s really hurting you. And do you want to pass along lab results to all of your doctors? You know, often you say yes,&lt;br /&gt;
 but there&#039;s some tests that you would only do if you had particular diseases. So, the fact that you had that test makes me, if I&#039;m a doctor, realize that you&#039;re probably HIV positive for something, you know, just to pick one. Um, so you might not realize the implications of letting something out or keeping something. I have a question for you. Um, so how do we make money?&lt;br /&gt;
 The good news is we&#039;re academics.&lt;br /&gt;
 &lt;br /&gt;
We&#039;re the guys who who figured out that we don&#039;t have to make money out of it. We took a vow of poverty when but I didn&#039;t take as much of a vow. So I like to be on boards that help companies.&lt;br /&gt;
 Yeah. So so I mean I was on a board of a startup that last week went bankrupt. You know could VC pulled out. You know I&#039;m happy to say the reason we lost it was because another company figured out how to do the semantic stuff better, faster, cheaper than we did. So you know I think I think the real answer is that it&#039;s It&#039;s an infrastructure for innovation. This data stuff in particular, you know, the government is saying to people, please find ways to make money off of this, right? The Democrats are saying, &amp;quot;Please find ways to make money because then the Republicans won&#039;t be able to turn it off if they if they become the winners.&amp;quot; The Republicans are saying, &amp;quot;Please make money because then we&#039;re able to show that, you know, we back the commercial part of this stuff, not the area or you know, that that crazy science crap, right?&amp;quot; So I mean sorry my political may show but again more seriously I mean you know so it it&#039;s really interesting across the aisle you&#039;re seeing different reasons they want to share this stuff but it all comes down to creating value via innovation or also saving money. So I thought Gail had a nice example where she said Foya it&#039;s a cost savings perspective. Um, so you know, it&#039;s a new way of looking at something that they were mandated to do anyway.&lt;br /&gt;
 Finding you what you&#039;re looking for when you don&#039;t know what you&#039;re looking for is the big value stuff. So, you know,&lt;br /&gt;
 again, we didn&#039;t talk a lot about the commercial edge of this stuff, the RDFA stuff that&#039;s going on now, things like that. I just sort of assume you&#039;ll have&lt;br /&gt;
 other people to mash up who talk more in that space. But, you know, I mean, 10 years ago, we said there&#039;s value in&lt;br /&gt;
 doing this semantic stuff and people kind of giggled a lot, right? And now 10 years later, we&#039;re not having a lot of trouble convincing people that there&lt;br /&gt;
 there&#039;s value somewhere in here. And we&#039;re being asked exactly the question,&lt;br /&gt;
 where is that? Where&#039;s the specific value? And that, you know, we didn&#039;t know how to answer that for the web for the first 15 years and turned out to be&lt;br /&gt;
 advertising instead of pornography. It surprised a lot of us. Um, you know,&lt;br /&gt;
 some of us thought it would be something better. Before we take the next question, someone mentioned seen before. You know, just take a look at Zene.&lt;br /&gt;
 haven&#039;t been there, you should check it out maybe of it. semantic technologies.&lt;br /&gt;
 It&#039;s the is the business of commercial leveraging the semantic web and most of us sort of started going there five or six years ago when you know the small&lt;br /&gt;
 amounts of people and people struggling to to sell the you know the snack web and suddenly Oracle shows up and all&lt;br /&gt;
 these startup companies are coming along and you know there&#039;s value in apps and again you know the the ad economy of&lt;br /&gt;
 the internet is driven by knowledge and what we&#039;re talking about is making some more of that explicit and machine uh&lt;br /&gt;
 manipulable. So you know the where the money is is easy. The how to extract that and commercialize that&#039;s the hard&lt;br /&gt;
 one. But you know search engine optimizers are looking are glee are looking glee at RDFA. They&#039;re like how&lt;br /&gt;
 do we get this everywhere? things like that where you know a year ago they were like that stuff will never catch on. We&#039;re seeing really amazing changes&lt;br /&gt;
 in people&#039;s thinking, but you know, I don&#039;t think it&#039;s yet hit the point. Did you have a question?&lt;br /&gt;
 Well, so I think actually the Alzheimer&#039;s is the most recent one that they said they couldn&#039;t have made the the findings without sharing that data.&lt;br /&gt;
 The sky surveys are some nice examples. When they shared dat a, you&#039;ve got high school teachers collaborating with astronomers because&lt;br /&gt;
 they one person saw in the data and it turned out to be something really of interest. I think actually we can point&lt;br /&gt;
 to you know at least dozens and hundreds now of examples where sharing the data and working in some kind of collaborative manner is starting to&lt;br /&gt;
 change the way discoveries are made. I Peter and I started this what a decade ago and I think we&#039;re starting to see um&lt;br /&gt;
 the our users changing the way they look at science. I remember Peter introduced me to one of his collaborators who was&lt;br /&gt;
 very hostile to this new way of doing science and he was like I don&#039;t need that you know I know how to do this and&lt;br /&gt;
 then when he decided off through his own intuition that this would change the&lt;br /&gt;
 way he looked at forming an experiment then he became our biggest supporter and I think that yeah speaking of Right.&lt;br /&gt;
&lt;br /&gt;
There&#039;s so with inside with inside FBI so we have become somewhat popular inside FBI as well and the single largest outreach to us is coming from the humanities social sciences cultural anthropologists because digital humanities is if you think science has taken off digital humanities is going through the roof computational economics computational sociology you know we&#039;re talking to people who are running things like asthma portals who are finding that the environmental factors is a real now back as a big thing and guess what you need environmental scientists and they don&#039;t speak the same language. So this is all over the place.&lt;br /&gt;
 I was on a phone call with the sort of chief visionary technologist guy at a at a medical company today and he said he had just come back from a meeting held between sort of some of the biggest companies in the world of the web and some of the biggest medical guys talking about the future. And he said, you know, the one thing everyone agreed on was 20 years from now, the hospital won&#039;t exist the way we think of it now, right? That it&#039;ll be sort of it&#039;s going through what libraries are going through now.&lt;br /&gt;
&lt;br /&gt;
That used to be the primary place you went to get information. Now it&#039;s not. And they have to reinvent, you know, what is their role? And they believe the hospital will stop being the primary place you go to get health care and health information and and to do monitoring and to do testing. Right? He said there were 20 different visions of what would replace it and how it would happen. But you know it&#039;s a big area an exciting area right now and you know the amazing thing to me is this technology is one of the things that people are agreeing is probably a big it it&#039;s not sufficient but it&#039;s necessary and you see a lot you see it at all levels both at the PhD researcher but also like patients like me have you heard of that website or trial. Yeah. So people who have some medical problem are sharing information about that medical problem and then they&#039;re forming communities. You know you can debate whether it&#039;s good or bad but it&#039;s you know a new phenomena that was enabled by this kind of technology. Just a quick follow up point on that. One interesting aspect of that is people are looking at their like patients like me and things like that where they&#039;re getting and they&#039;re self tagging like stage bio as well. They&#039;re selftagging it with semantic information that&#039;s shared among the community and then that information is being funneled back to their scientists for the phenotype information. So then they correlate people&#039;s actual you know what they&#039;re what they&#039;re experiencing in their life in terms of symptoms or or whatnot back to the genetic information. so it&#039;s an interesting case of kind of like spec data flowing in in multiple directions where the the genotype data is semantic and people take phenotype data that&#039;s essentially mashed up by organizations like patients like me and then then it&#039;s new it&#039;s very interesting as well and just to conclude this the question asked before are the vested interests who control this stuff you know hostile  like you wouldn&#039;t believe that ownership of information is the biggest thing in medicine today. The sharing of information is the biggest thing in healthcare tomorrow. Talk about talk about, you know, a conflict that our society will have to face sooner or later in a very big way. And and I mean, yeah, it&#039;s fun to be one of the technologists in this space. Boy, I keep my head down a lot when we get into those conversations. Absolutely. I think that&#039;s probably a good time.&lt;br /&gt;
 &lt;br /&gt;
Marco, that&#039;s probably a good time to stop because you wanted a few minutes. Thank you so much for coming.&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6638</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6638"/>
		<updated>2026-04-16T21:25:41Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;&amp;lt;youtube&amp;gt;https://youtu.be/SWVO-Pviz1w&amp;lt;/youtube&amp;gt;&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
[https://en.wikipedia.org/wiki/Peter_Fox_(professor) http://www.rpi.edu/dept/ees/people/faculty/fox.html]&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Gale A. Brewer is the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
===Resources===&lt;br /&gt;
[[Data Gov Transcript]]&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6637</id>
		<title>Meetup at the Semantic Technology Conference 2010 in San Francisco</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6637"/>
		<updated>2026-04-04T08:54:21Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__NOTOC__&lt;br /&gt;
&lt;br /&gt;
This year our global Lotico Semantic Web event took place at the Semantic Technology Conference 2010 in San Francisco. Come and meet your peers to discuss all things Semantic Web, Web 3.0 and Linked Data to make the Web of Data a reality. It took more than 10 years to get the Semantic Web initiative where it is today and we have good reason to believe that it&#039;s about time to hit the mainstream web. This I believe will not happen without friction since the standards in the Semantic Web initiative are geared towards a more academic audience rather than web practitioners. With the growing adoption in the mainstream community this might require some minor review of some recommendations, additional training documentation and more tools for development. So it&#039;s interesting times again, I hope to see you in San Francisco this summer.&lt;br /&gt;
&lt;br /&gt;
==Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;&amp;lt;youtube width=&amp;quot;800&amp;quot; height=600&amp;gt;https://youtu.be/vATVkMTv8q8&amp;lt;/youtube&amp;gt;&amp;lt;center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[http://vimeo.com/20316782 Vimeo Video]&lt;br /&gt;
&lt;br /&gt;
Smart mobs emerge when communication and computing technologies amplify human talents for cooperation. The impacts of smart mob technology already appear to be both beneficial and destructive, used by some of its earliest adopters to support democracy and by others to coordinate terrorist attacks. The technologies that are beginning to make smart mobs possible are mobile communication devices and pervasive computing - inexpensive microprocessors embedded in everyday objects and environments. Already, governments have fallen, youth subcultures have blossomed from Asia to Scandinavia, new industries have been born and older industries have launched furious counterattacks.&lt;br /&gt;
&lt;br /&gt;
Street demonstrators in the 1999 anti-WTO protests used dynamically updated websites, cell-phones, and &amp;quot;swarming&amp;quot; tactics in the &amp;quot;battle of Seattle.&amp;quot; A million Filipinos toppled President Estrada through public demonstrations organized through salvos of text messages.&lt;br /&gt;
&lt;br /&gt;
The pieces of the puzzle are all around us now, but haven&#039;t joined together yet. The radio chips designed to replace barcodes on manufactured objects are part of it. Wireless Internet nodes in cafes, hotels, and neighborhoods are part of it. Millions of people who lend their computers to the search for extraterrestrial intelligence are part of it. The way buyers and sellers rate each other on Internet auction site eBay is part of it. Research by biologists, sociologists, and economists into the nature of cooperation offer explanatory frameworks. At least one key global business question is part of it - why is the Japanese company DoCoMo profiting from enhanced wireless Internet services while US and European mobile telephony operators struggle to avoid failure?&lt;br /&gt;
&lt;br /&gt;
The people who make up smart mobs cooperate in ways never before possible because they carry devices that possess both communication and computing capabilities. Their mobile devices connect them with other information devices in the environment as well as with other people&#039;s telephones. Dirt-cheap microprocessors embedded in everything from box tops to shoes are beginning to permeate furniture, buildings, neighborhoods, products with invisible intercommunicating smartifacts. When they connect the tangible objects and places of our daily lives with the Internet, handheld communication media mutate into wearable remote control devices for the physical world.&lt;br /&gt;
&lt;br /&gt;
Media cartels and government agencies are seeking to reimpose the regime of the broadcast era in which the customers of technology will be deprived of the power to create and left only with the power to consume. That power struggle is what the battles over file-sharing, copy-protection, regulation of the radio spectrum are about. Are the populations of tomorrow going to be users, like the PC owners and website creators who turned technology to widespread innovation? Or will they be consumers, constrained from innovation and locked into the technology and business models of the most powerful entrenched interests?&lt;br /&gt;
&lt;br /&gt;
Howard Rheingold is one of the world&#039;s foremost authorities on the social implications of technology. Over the past twenty years he has traveled around the world, observing and writing about emerging trends in computing, communications, and culture. One of the creators and former founding executive editor of HotWired, he has served as editor of The Whole Earth Review, editor-in-chief of The Millennium Whole Earth Catalog, and on-line host for The Well. The author of several books, including The Virtual Community, Virtual Reality, and Tools for Thought, he lives in Mill Valley, California.&lt;br /&gt;
&lt;br /&gt;
== Participating Meetup groups (5093 members) ==&lt;br /&gt;
&lt;br /&gt;
[[Atlanta Semantic Web Meetup|Atlanta]]  -  [[Austin Semantic Web Meetup|Austin]]  - [[Lotico CSW Berlin|Berlin]] -  [[Cambridge Semantic Web Meetup|Cambridge]]  -  [[The Chicago Semantic Web Meetup Group|Chicago]] - [[Frankfurt Semantic Web Meetup|Frankfurt]]  - [[Central Florida Semantic Web Meetup|Central Florida]] -  [[London Semantic Web Meetup|London]]  -  [[Los Angeles Semantic Web Meetup|Los Angeles]]  - [[Munchen-Semantic-Web-Meetup | Munich]] - [[SWNYC|New York]]  -  [[Philadelphia Semantic Web Meetup|Philadelphia]]  - [[Oslo Semantic Web Meetup|Oslo]]  - [[Ottawa Semantic Web Meetup|Ottawa]] -  [[Princeton Semantic Web Meetup|Princeton]]  -  [[San Diego Semantic Web Meetup|San Diego]]  -  [[San Francisco Semantic Web Meetup|San Francisco]]  - [[Santiago Semantic Web Meetup|Santiago]] - [[Seattle Semantic Web Meetup|Seattle]]  -  [[Silicon Valley Meetup|Silicon Valley]]  -  [[Thessaloniki Semantic Web Meetup|Thessaloniki]]  -  [[Toronto Semantic Web Meetup|Toronto]]  -  [[Vancouver Semantic Web Meetup|Vancouver]]  -  [[Vienna Semantic Web Meetup|Vienna]]  -  [[Washington Semantic Web Meetup|Washington DC]]&lt;br /&gt;
&lt;br /&gt;
== Proposed Sessions ==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TABLE CELLSPACING=3 CELLPADDING=3&amp;gt;&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Monday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 21&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#EEEEEE&amp;gt;Pre-conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt; &amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Tuesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 22&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;First Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Global Semantic Web Meetup&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Wednesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 23&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Second Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[[Semantic Code Camp 2010|Semantic Unconference]]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Thursday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 24&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Third Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Semantic Social Networks - Meetup Ontology 2pm-3pm Incentives &amp;amp; Roadblocks for Participating in the Semantic Web 7pm&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Friday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 25&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Fourth Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[http://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/calendar/13126881/ LODE: Linking Open Descriptions of Events]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/TABLE&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Proposed Session Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Incentives &amp;amp; Roadblocks for Participating in the Semantic Web]]&#039;&#039;&#039; - Griffin Caprio&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Semantic Code Camp 2010]]&#039;&#039;&#039; - Shamod Lacoul&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Big Meetup Social 2010|Global Meetup Social 2010]]&#039;&#039;&#039; - Marco Neumann&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6636</id>
		<title>Meetup at the Semantic Technology Conference 2010 in San Francisco</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6636"/>
		<updated>2026-04-04T08:54:01Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__NOTOC__&lt;br /&gt;
&lt;br /&gt;
This year our global Lotico Semantic Web event took place at the Semantic Technology Conference 2010 in San Francisco. Come and meet your peers to discuss all things Semantic Web, Web 3.0 and Linked Data to make the Web of Data a reality. It took more than 10 years to get the Semantic Web initiative where it is today and we have good reason to believe that it&#039;s about time to hit the mainstream web. This I believe will not happen without friction since the standards in the Semantic Web initiative are geared towards a more academic audience rather than web practitioners. With the growing adoption in the mainstream community this might require some minor review of some recommendations, additional training documentation and more tools for development. So it&#039;s interesting times again, I hope to see you in San Francisco this summer.&lt;br /&gt;
&lt;br /&gt;
==Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;800&amp;quot; height=600&amp;gt;https://youtu.be/vATVkMTv8q8&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[http://vimeo.com/20316782 Vimeo Video]&lt;br /&gt;
&lt;br /&gt;
Smart mobs emerge when communication and computing technologies amplify human talents for cooperation. The impacts of smart mob technology already appear to be both beneficial and destructive, used by some of its earliest adopters to support democracy and by others to coordinate terrorist attacks. The technologies that are beginning to make smart mobs possible are mobile communication devices and pervasive computing - inexpensive microprocessors embedded in everyday objects and environments. Already, governments have fallen, youth subcultures have blossomed from Asia to Scandinavia, new industries have been born and older industries have launched furious counterattacks.&lt;br /&gt;
&lt;br /&gt;
Street demonstrators in the 1999 anti-WTO protests used dynamically updated websites, cell-phones, and &amp;quot;swarming&amp;quot; tactics in the &amp;quot;battle of Seattle.&amp;quot; A million Filipinos toppled President Estrada through public demonstrations organized through salvos of text messages.&lt;br /&gt;
&lt;br /&gt;
The pieces of the puzzle are all around us now, but haven&#039;t joined together yet. The radio chips designed to replace barcodes on manufactured objects are part of it. Wireless Internet nodes in cafes, hotels, and neighborhoods are part of it. Millions of people who lend their computers to the search for extraterrestrial intelligence are part of it. The way buyers and sellers rate each other on Internet auction site eBay is part of it. Research by biologists, sociologists, and economists into the nature of cooperation offer explanatory frameworks. At least one key global business question is part of it - why is the Japanese company DoCoMo profiting from enhanced wireless Internet services while US and European mobile telephony operators struggle to avoid failure?&lt;br /&gt;
&lt;br /&gt;
The people who make up smart mobs cooperate in ways never before possible because they carry devices that possess both communication and computing capabilities. Their mobile devices connect them with other information devices in the environment as well as with other people&#039;s telephones. Dirt-cheap microprocessors embedded in everything from box tops to shoes are beginning to permeate furniture, buildings, neighborhoods, products with invisible intercommunicating smartifacts. When they connect the tangible objects and places of our daily lives with the Internet, handheld communication media mutate into wearable remote control devices for the physical world.&lt;br /&gt;
&lt;br /&gt;
Media cartels and government agencies are seeking to reimpose the regime of the broadcast era in which the customers of technology will be deprived of the power to create and left only with the power to consume. That power struggle is what the battles over file-sharing, copy-protection, regulation of the radio spectrum are about. Are the populations of tomorrow going to be users, like the PC owners and website creators who turned technology to widespread innovation? Or will they be consumers, constrained from innovation and locked into the technology and business models of the most powerful entrenched interests?&lt;br /&gt;
&lt;br /&gt;
Howard Rheingold is one of the world&#039;s foremost authorities on the social implications of technology. Over the past twenty years he has traveled around the world, observing and writing about emerging trends in computing, communications, and culture. One of the creators and former founding executive editor of HotWired, he has served as editor of The Whole Earth Review, editor-in-chief of The Millennium Whole Earth Catalog, and on-line host for The Well. The author of several books, including The Virtual Community, Virtual Reality, and Tools for Thought, he lives in Mill Valley, California.&lt;br /&gt;
&lt;br /&gt;
== Participating Meetup groups (5093 members) ==&lt;br /&gt;
&lt;br /&gt;
[[Atlanta Semantic Web Meetup|Atlanta]]  -  [[Austin Semantic Web Meetup|Austin]]  - [[Lotico CSW Berlin|Berlin]] -  [[Cambridge Semantic Web Meetup|Cambridge]]  -  [[The Chicago Semantic Web Meetup Group|Chicago]] - [[Frankfurt Semantic Web Meetup|Frankfurt]]  - [[Central Florida Semantic Web Meetup|Central Florida]] -  [[London Semantic Web Meetup|London]]  -  [[Los Angeles Semantic Web Meetup|Los Angeles]]  - [[Munchen-Semantic-Web-Meetup | Munich]] - [[SWNYC|New York]]  -  [[Philadelphia Semantic Web Meetup|Philadelphia]]  - [[Oslo Semantic Web Meetup|Oslo]]  - [[Ottawa Semantic Web Meetup|Ottawa]] -  [[Princeton Semantic Web Meetup|Princeton]]  -  [[San Diego Semantic Web Meetup|San Diego]]  -  [[San Francisco Semantic Web Meetup|San Francisco]]  - [[Santiago Semantic Web Meetup|Santiago]] - [[Seattle Semantic Web Meetup|Seattle]]  -  [[Silicon Valley Meetup|Silicon Valley]]  -  [[Thessaloniki Semantic Web Meetup|Thessaloniki]]  -  [[Toronto Semantic Web Meetup|Toronto]]  -  [[Vancouver Semantic Web Meetup|Vancouver]]  -  [[Vienna Semantic Web Meetup|Vienna]]  -  [[Washington Semantic Web Meetup|Washington DC]]&lt;br /&gt;
&lt;br /&gt;
== Proposed Sessions ==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TABLE CELLSPACING=3 CELLPADDING=3&amp;gt;&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Monday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 21&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#EEEEEE&amp;gt;Pre-conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt; &amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Tuesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 22&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;First Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Global Semantic Web Meetup&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Wednesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 23&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Second Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[[Semantic Code Camp 2010|Semantic Unconference]]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Thursday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 24&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Third Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Semantic Social Networks - Meetup Ontology 2pm-3pm Incentives &amp;amp; Roadblocks for Participating in the Semantic Web 7pm&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Friday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 25&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Fourth Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[http://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/calendar/13126881/ LODE: Linking Open Descriptions of Events]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/TABLE&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Proposed Session Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Incentives &amp;amp; Roadblocks for Participating in the Semantic Web]]&#039;&#039;&#039; - Griffin Caprio&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Semantic Code Camp 2010]]&#039;&#039;&#039; - Shamod Lacoul&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Big Meetup Social 2010|Global Meetup Social 2010]]&#039;&#039;&#039; - Marco Neumann&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6635</id>
		<title>Meetup at the Semantic Technology Conference 2010 in San Francisco</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6635"/>
		<updated>2026-04-04T08:53:39Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__NOTOC__&lt;br /&gt;
&lt;br /&gt;
This year our global Lotico Semantic Web event took place at the Semantic Technology Conference 2010 in San Francisco. Come and meet your peers to discuss all things Semantic Web, Web 3.0 and Linked Data to make the Web of Data a reality. It took more than 10 years to get the Semantic Web initiative where it is today and we have good reason to believe that it&#039;s about time to hit the mainstream web. This I believe will not happen without friction since the standards in the Semantic Web initiative are geared towards a more academic audience rather than web practitioners. With the growing adoption in the mainstream community this might require some minor review of some recommendations, additional training documentation and more tools for development. So it&#039;s interesting times again, I hope to see you in San Francisco this summer.&lt;br /&gt;
&lt;br /&gt;
==Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;800&amp;quot;&amp;gt;https://youtu.be/vATVkMTv8q8&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[http://vimeo.com/20316782 Vimeo Video]&lt;br /&gt;
&lt;br /&gt;
Smart mobs emerge when communication and computing technologies amplify human talents for cooperation. The impacts of smart mob technology already appear to be both beneficial and destructive, used by some of its earliest adopters to support democracy and by others to coordinate terrorist attacks. The technologies that are beginning to make smart mobs possible are mobile communication devices and pervasive computing - inexpensive microprocessors embedded in everyday objects and environments. Already, governments have fallen, youth subcultures have blossomed from Asia to Scandinavia, new industries have been born and older industries have launched furious counterattacks.&lt;br /&gt;
&lt;br /&gt;
Street demonstrators in the 1999 anti-WTO protests used dynamically updated websites, cell-phones, and &amp;quot;swarming&amp;quot; tactics in the &amp;quot;battle of Seattle.&amp;quot; A million Filipinos toppled President Estrada through public demonstrations organized through salvos of text messages.&lt;br /&gt;
&lt;br /&gt;
The pieces of the puzzle are all around us now, but haven&#039;t joined together yet. The radio chips designed to replace barcodes on manufactured objects are part of it. Wireless Internet nodes in cafes, hotels, and neighborhoods are part of it. Millions of people who lend their computers to the search for extraterrestrial intelligence are part of it. The way buyers and sellers rate each other on Internet auction site eBay is part of it. Research by biologists, sociologists, and economists into the nature of cooperation offer explanatory frameworks. At least one key global business question is part of it - why is the Japanese company DoCoMo profiting from enhanced wireless Internet services while US and European mobile telephony operators struggle to avoid failure?&lt;br /&gt;
&lt;br /&gt;
The people who make up smart mobs cooperate in ways never before possible because they carry devices that possess both communication and computing capabilities. Their mobile devices connect them with other information devices in the environment as well as with other people&#039;s telephones. Dirt-cheap microprocessors embedded in everything from box tops to shoes are beginning to permeate furniture, buildings, neighborhoods, products with invisible intercommunicating smartifacts. When they connect the tangible objects and places of our daily lives with the Internet, handheld communication media mutate into wearable remote control devices for the physical world.&lt;br /&gt;
&lt;br /&gt;
Media cartels and government agencies are seeking to reimpose the regime of the broadcast era in which the customers of technology will be deprived of the power to create and left only with the power to consume. That power struggle is what the battles over file-sharing, copy-protection, regulation of the radio spectrum are about. Are the populations of tomorrow going to be users, like the PC owners and website creators who turned technology to widespread innovation? Or will they be consumers, constrained from innovation and locked into the technology and business models of the most powerful entrenched interests?&lt;br /&gt;
&lt;br /&gt;
Howard Rheingold is one of the world&#039;s foremost authorities on the social implications of technology. Over the past twenty years he has traveled around the world, observing and writing about emerging trends in computing, communications, and culture. One of the creators and former founding executive editor of HotWired, he has served as editor of The Whole Earth Review, editor-in-chief of The Millennium Whole Earth Catalog, and on-line host for The Well. The author of several books, including The Virtual Community, Virtual Reality, and Tools for Thought, he lives in Mill Valley, California.&lt;br /&gt;
&lt;br /&gt;
== Participating Meetup groups (5093 members) ==&lt;br /&gt;
&lt;br /&gt;
[[Atlanta Semantic Web Meetup|Atlanta]]  -  [[Austin Semantic Web Meetup|Austin]]  - [[Lotico CSW Berlin|Berlin]] -  [[Cambridge Semantic Web Meetup|Cambridge]]  -  [[The Chicago Semantic Web Meetup Group|Chicago]] - [[Frankfurt Semantic Web Meetup|Frankfurt]]  - [[Central Florida Semantic Web Meetup|Central Florida]] -  [[London Semantic Web Meetup|London]]  -  [[Los Angeles Semantic Web Meetup|Los Angeles]]  - [[Munchen-Semantic-Web-Meetup | Munich]] - [[SWNYC|New York]]  -  [[Philadelphia Semantic Web Meetup|Philadelphia]]  - [[Oslo Semantic Web Meetup|Oslo]]  - [[Ottawa Semantic Web Meetup|Ottawa]] -  [[Princeton Semantic Web Meetup|Princeton]]  -  [[San Diego Semantic Web Meetup|San Diego]]  -  [[San Francisco Semantic Web Meetup|San Francisco]]  - [[Santiago Semantic Web Meetup|Santiago]] - [[Seattle Semantic Web Meetup|Seattle]]  -  [[Silicon Valley Meetup|Silicon Valley]]  -  [[Thessaloniki Semantic Web Meetup|Thessaloniki]]  -  [[Toronto Semantic Web Meetup|Toronto]]  -  [[Vancouver Semantic Web Meetup|Vancouver]]  -  [[Vienna Semantic Web Meetup|Vienna]]  -  [[Washington Semantic Web Meetup|Washington DC]]&lt;br /&gt;
&lt;br /&gt;
== Proposed Sessions ==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TABLE CELLSPACING=3 CELLPADDING=3&amp;gt;&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Monday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 21&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#EEEEEE&amp;gt;Pre-conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt; &amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Tuesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 22&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;First Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Global Semantic Web Meetup&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Wednesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 23&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Second Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[[Semantic Code Camp 2010|Semantic Unconference]]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Thursday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 24&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Third Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Semantic Social Networks - Meetup Ontology 2pm-3pm Incentives &amp;amp; Roadblocks for Participating in the Semantic Web 7pm&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Friday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 25&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Fourth Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[http://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/calendar/13126881/ LODE: Linking Open Descriptions of Events]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/TABLE&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Proposed Session Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Incentives &amp;amp; Roadblocks for Participating in the Semantic Web]]&#039;&#039;&#039; - Griffin Caprio&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Semantic Code Camp 2010]]&#039;&#039;&#039; - Shamod Lacoul&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Big Meetup Social 2010|Global Meetup Social 2010]]&#039;&#039;&#039; - Marco Neumann&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6634</id>
		<title>Meetup at the Semantic Technology Conference 2010 in San Francisco</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6634"/>
		<updated>2026-04-04T08:53:30Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__NOTOC__&lt;br /&gt;
&lt;br /&gt;
This year our global Lotico Semantic Web event took place at the Semantic Technology Conference 2010 in San Francisco. Come and meet your peers to discuss all things Semantic Web, Web 3.0 and Linked Data to make the Web of Data a reality. It took more than 10 years to get the Semantic Web initiative where it is today and we have good reason to believe that it&#039;s about time to hit the mainstream web. This I believe will not happen without friction since the standards in the Semantic Web initiative are geared towards a more academic audience rather than web practitioners. With the growing adoption in the mainstream community this might require some minor review of some recommendations, additional training documentation and more tools for development. So it&#039;s interesting times again, I hope to see you in San Francisco this summer.&lt;br /&gt;
&lt;br /&gt;
==Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;400&amp;quot;&amp;gt;https://youtu.be/vATVkMTv8q8&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[http://vimeo.com/20316782 Vimeo Video]&lt;br /&gt;
&lt;br /&gt;
Smart mobs emerge when communication and computing technologies amplify human talents for cooperation. The impacts of smart mob technology already appear to be both beneficial and destructive, used by some of its earliest adopters to support democracy and by others to coordinate terrorist attacks. The technologies that are beginning to make smart mobs possible are mobile communication devices and pervasive computing - inexpensive microprocessors embedded in everyday objects and environments. Already, governments have fallen, youth subcultures have blossomed from Asia to Scandinavia, new industries have been born and older industries have launched furious counterattacks.&lt;br /&gt;
&lt;br /&gt;
Street demonstrators in the 1999 anti-WTO protests used dynamically updated websites, cell-phones, and &amp;quot;swarming&amp;quot; tactics in the &amp;quot;battle of Seattle.&amp;quot; A million Filipinos toppled President Estrada through public demonstrations organized through salvos of text messages.&lt;br /&gt;
&lt;br /&gt;
The pieces of the puzzle are all around us now, but haven&#039;t joined together yet. The radio chips designed to replace barcodes on manufactured objects are part of it. Wireless Internet nodes in cafes, hotels, and neighborhoods are part of it. Millions of people who lend their computers to the search for extraterrestrial intelligence are part of it. The way buyers and sellers rate each other on Internet auction site eBay is part of it. Research by biologists, sociologists, and economists into the nature of cooperation offer explanatory frameworks. At least one key global business question is part of it - why is the Japanese company DoCoMo profiting from enhanced wireless Internet services while US and European mobile telephony operators struggle to avoid failure?&lt;br /&gt;
&lt;br /&gt;
The people who make up smart mobs cooperate in ways never before possible because they carry devices that possess both communication and computing capabilities. Their mobile devices connect them with other information devices in the environment as well as with other people&#039;s telephones. Dirt-cheap microprocessors embedded in everything from box tops to shoes are beginning to permeate furniture, buildings, neighborhoods, products with invisible intercommunicating smartifacts. When they connect the tangible objects and places of our daily lives with the Internet, handheld communication media mutate into wearable remote control devices for the physical world.&lt;br /&gt;
&lt;br /&gt;
Media cartels and government agencies are seeking to reimpose the regime of the broadcast era in which the customers of technology will be deprived of the power to create and left only with the power to consume. That power struggle is what the battles over file-sharing, copy-protection, regulation of the radio spectrum are about. Are the populations of tomorrow going to be users, like the PC owners and website creators who turned technology to widespread innovation? Or will they be consumers, constrained from innovation and locked into the technology and business models of the most powerful entrenched interests?&lt;br /&gt;
&lt;br /&gt;
Howard Rheingold is one of the world&#039;s foremost authorities on the social implications of technology. Over the past twenty years he has traveled around the world, observing and writing about emerging trends in computing, communications, and culture. One of the creators and former founding executive editor of HotWired, he has served as editor of The Whole Earth Review, editor-in-chief of The Millennium Whole Earth Catalog, and on-line host for The Well. The author of several books, including The Virtual Community, Virtual Reality, and Tools for Thought, he lives in Mill Valley, California.&lt;br /&gt;
&lt;br /&gt;
== Participating Meetup groups (5093 members) ==&lt;br /&gt;
&lt;br /&gt;
[[Atlanta Semantic Web Meetup|Atlanta]]  -  [[Austin Semantic Web Meetup|Austin]]  - [[Lotico CSW Berlin|Berlin]] -  [[Cambridge Semantic Web Meetup|Cambridge]]  -  [[The Chicago Semantic Web Meetup Group|Chicago]] - [[Frankfurt Semantic Web Meetup|Frankfurt]]  - [[Central Florida Semantic Web Meetup|Central Florida]] -  [[London Semantic Web Meetup|London]]  -  [[Los Angeles Semantic Web Meetup|Los Angeles]]  - [[Munchen-Semantic-Web-Meetup | Munich]] - [[SWNYC|New York]]  -  [[Philadelphia Semantic Web Meetup|Philadelphia]]  - [[Oslo Semantic Web Meetup|Oslo]]  - [[Ottawa Semantic Web Meetup|Ottawa]] -  [[Princeton Semantic Web Meetup|Princeton]]  -  [[San Diego Semantic Web Meetup|San Diego]]  -  [[San Francisco Semantic Web Meetup|San Francisco]]  - [[Santiago Semantic Web Meetup|Santiago]] - [[Seattle Semantic Web Meetup|Seattle]]  -  [[Silicon Valley Meetup|Silicon Valley]]  -  [[Thessaloniki Semantic Web Meetup|Thessaloniki]]  -  [[Toronto Semantic Web Meetup|Toronto]]  -  [[Vancouver Semantic Web Meetup|Vancouver]]  -  [[Vienna Semantic Web Meetup|Vienna]]  -  [[Washington Semantic Web Meetup|Washington DC]]&lt;br /&gt;
&lt;br /&gt;
== Proposed Sessions ==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TABLE CELLSPACING=3 CELLPADDING=3&amp;gt;&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Monday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 21&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#EEEEEE&amp;gt;Pre-conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt; &amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Tuesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 22&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;First Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Global Semantic Web Meetup&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Wednesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 23&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Second Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[[Semantic Code Camp 2010|Semantic Unconference]]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Thursday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 24&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Third Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Semantic Social Networks - Meetup Ontology 2pm-3pm Incentives &amp;amp; Roadblocks for Participating in the Semantic Web 7pm&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Friday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 25&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Fourth Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[http://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/calendar/13126881/ LODE: Linking Open Descriptions of Events]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/TABLE&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Proposed Session Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Incentives &amp;amp; Roadblocks for Participating in the Semantic Web]]&#039;&#039;&#039; - Griffin Caprio&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Semantic Code Camp 2010]]&#039;&#039;&#039; - Shamod Lacoul&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Big Meetup Social 2010|Global Meetup Social 2010]]&#039;&#039;&#039; - Marco Neumann&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6633</id>
		<title>Meetup at the Semantic Technology Conference 2010 in San Francisco</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6633"/>
		<updated>2026-04-04T08:52:56Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__NOTOC__&lt;br /&gt;
&lt;br /&gt;
This year our global Lotico Semantic Web event took place at the Semantic Technology Conference 2010 in San Francisco. Come and meet your peers to discuss all things Semantic Web, Web 3.0 and Linked Data to make the Web of Data a reality. It took more than 10 years to get the Semantic Web initiative where it is today and we have good reason to believe that it&#039;s about time to hit the mainstream web. This I believe will not happen without friction since the standards in the Semantic Web initiative are geared towards a more academic audience rather than web practitioners. With the growing adoption in the mainstream community this might require some minor review of some recommendations, additional training documentation and more tools for development. So it&#039;s interesting times again, I hope to see you in San Francisco this summer.&lt;br /&gt;
&lt;br /&gt;
==Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;youtube width=&amp;quot;80%&amp;quot;&amp;gt;https://youtu.be/vATVkMTv8q8&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[http://vimeo.com/20316782 Vimeo Video]&lt;br /&gt;
&lt;br /&gt;
Smart mobs emerge when communication and computing technologies amplify human talents for cooperation. The impacts of smart mob technology already appear to be both beneficial and destructive, used by some of its earliest adopters to support democracy and by others to coordinate terrorist attacks. The technologies that are beginning to make smart mobs possible are mobile communication devices and pervasive computing - inexpensive microprocessors embedded in everyday objects and environments. Already, governments have fallen, youth subcultures have blossomed from Asia to Scandinavia, new industries have been born and older industries have launched furious counterattacks.&lt;br /&gt;
&lt;br /&gt;
Street demonstrators in the 1999 anti-WTO protests used dynamically updated websites, cell-phones, and &amp;quot;swarming&amp;quot; tactics in the &amp;quot;battle of Seattle.&amp;quot; A million Filipinos toppled President Estrada through public demonstrations organized through salvos of text messages.&lt;br /&gt;
&lt;br /&gt;
The pieces of the puzzle are all around us now, but haven&#039;t joined together yet. The radio chips designed to replace barcodes on manufactured objects are part of it. Wireless Internet nodes in cafes, hotels, and neighborhoods are part of it. Millions of people who lend their computers to the search for extraterrestrial intelligence are part of it. The way buyers and sellers rate each other on Internet auction site eBay is part of it. Research by biologists, sociologists, and economists into the nature of cooperation offer explanatory frameworks. At least one key global business question is part of it - why is the Japanese company DoCoMo profiting from enhanced wireless Internet services while US and European mobile telephony operators struggle to avoid failure?&lt;br /&gt;
&lt;br /&gt;
The people who make up smart mobs cooperate in ways never before possible because they carry devices that possess both communication and computing capabilities. Their mobile devices connect them with other information devices in the environment as well as with other people&#039;s telephones. Dirt-cheap microprocessors embedded in everything from box tops to shoes are beginning to permeate furniture, buildings, neighborhoods, products with invisible intercommunicating smartifacts. When they connect the tangible objects and places of our daily lives with the Internet, handheld communication media mutate into wearable remote control devices for the physical world.&lt;br /&gt;
&lt;br /&gt;
Media cartels and government agencies are seeking to reimpose the regime of the broadcast era in which the customers of technology will be deprived of the power to create and left only with the power to consume. That power struggle is what the battles over file-sharing, copy-protection, regulation of the radio spectrum are about. Are the populations of tomorrow going to be users, like the PC owners and website creators who turned technology to widespread innovation? Or will they be consumers, constrained from innovation and locked into the technology and business models of the most powerful entrenched interests?&lt;br /&gt;
&lt;br /&gt;
Howard Rheingold is one of the world&#039;s foremost authorities on the social implications of technology. Over the past twenty years he has traveled around the world, observing and writing about emerging trends in computing, communications, and culture. One of the creators and former founding executive editor of HotWired, he has served as editor of The Whole Earth Review, editor-in-chief of The Millennium Whole Earth Catalog, and on-line host for The Well. The author of several books, including The Virtual Community, Virtual Reality, and Tools for Thought, he lives in Mill Valley, California.&lt;br /&gt;
&lt;br /&gt;
== Participating Meetup groups (5093 members) ==&lt;br /&gt;
&lt;br /&gt;
[[Atlanta Semantic Web Meetup|Atlanta]]  -  [[Austin Semantic Web Meetup|Austin]]  - [[Lotico CSW Berlin|Berlin]] -  [[Cambridge Semantic Web Meetup|Cambridge]]  -  [[The Chicago Semantic Web Meetup Group|Chicago]] - [[Frankfurt Semantic Web Meetup|Frankfurt]]  - [[Central Florida Semantic Web Meetup|Central Florida]] -  [[London Semantic Web Meetup|London]]  -  [[Los Angeles Semantic Web Meetup|Los Angeles]]  - [[Munchen-Semantic-Web-Meetup | Munich]] - [[SWNYC|New York]]  -  [[Philadelphia Semantic Web Meetup|Philadelphia]]  - [[Oslo Semantic Web Meetup|Oslo]]  - [[Ottawa Semantic Web Meetup|Ottawa]] -  [[Princeton Semantic Web Meetup|Princeton]]  -  [[San Diego Semantic Web Meetup|San Diego]]  -  [[San Francisco Semantic Web Meetup|San Francisco]]  - [[Santiago Semantic Web Meetup|Santiago]] - [[Seattle Semantic Web Meetup|Seattle]]  -  [[Silicon Valley Meetup|Silicon Valley]]  -  [[Thessaloniki Semantic Web Meetup|Thessaloniki]]  -  [[Toronto Semantic Web Meetup|Toronto]]  -  [[Vancouver Semantic Web Meetup|Vancouver]]  -  [[Vienna Semantic Web Meetup|Vienna]]  -  [[Washington Semantic Web Meetup|Washington DC]]&lt;br /&gt;
&lt;br /&gt;
== Proposed Sessions ==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TABLE CELLSPACING=3 CELLPADDING=3&amp;gt;&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Monday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 21&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#EEEEEE&amp;gt;Pre-conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt; &amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Tuesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 22&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;First Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Global Semantic Web Meetup&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Wednesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 23&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Second Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[[Semantic Code Camp 2010|Semantic Unconference]]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Thursday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 24&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Third Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Semantic Social Networks - Meetup Ontology 2pm-3pm Incentives &amp;amp; Roadblocks for Participating in the Semantic Web 7pm&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Friday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 25&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Fourth Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[http://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/calendar/13126881/ LODE: Linking Open Descriptions of Events]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/TABLE&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Proposed Session Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Incentives &amp;amp; Roadblocks for Participating in the Semantic Web]]&#039;&#039;&#039; - Griffin Caprio&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Semantic Code Camp 2010]]&#039;&#039;&#039; - Shamod Lacoul&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Big Meetup Social 2010|Global Meetup Social 2010]]&#039;&#039;&#039; - Marco Neumann&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6632</id>
		<title>Meetup at the Semantic Technology Conference 2010 in San Francisco</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6632"/>
		<updated>2026-04-04T08:44:50Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__NOTOC__&lt;br /&gt;
&lt;br /&gt;
This year our global Lotico Semantic Web event took place at the Semantic Technology Conference 2010 in San Francisco. Come and meet your peers to discuss all things Semantic Web, Web 3.0 and Linked Data to make the Web of Data a reality. It took more than 10 years to get the Semantic Web initiative where it is today and we have good reason to believe that it&#039;s about time to hit the mainstream web. This I believe will not happen without friction since the standards in the Semantic Web initiative are geared towards a more academic audience rather than web practitioners. With the growing adoption in the mainstream community this might require some minor review of some recommendations, additional training documentation and more tools for development. So it&#039;s interesting times again, I hope to see you in San Francisco this summer.&lt;br /&gt;
&lt;br /&gt;
==Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;youtube&amp;gt;https://youtu.be/vATVkMTv8q8&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[http://vimeo.com/20316782 Vimeo Video]&lt;br /&gt;
&lt;br /&gt;
Smart mobs emerge when communication and computing technologies amplify human talents for cooperation. The impacts of smart mob technology already appear to be both beneficial and destructive, used by some of its earliest adopters to support democracy and by others to coordinate terrorist attacks. The technologies that are beginning to make smart mobs possible are mobile communication devices and pervasive computing - inexpensive microprocessors embedded in everyday objects and environments. Already, governments have fallen, youth subcultures have blossomed from Asia to Scandinavia, new industries have been born and older industries have launched furious counterattacks.&lt;br /&gt;
&lt;br /&gt;
Street demonstrators in the 1999 anti-WTO protests used dynamically updated websites, cell-phones, and &amp;quot;swarming&amp;quot; tactics in the &amp;quot;battle of Seattle.&amp;quot; A million Filipinos toppled President Estrada through public demonstrations organized through salvos of text messages.&lt;br /&gt;
&lt;br /&gt;
The pieces of the puzzle are all around us now, but haven&#039;t joined together yet. The radio chips designed to replace barcodes on manufactured objects are part of it. Wireless Internet nodes in cafes, hotels, and neighborhoods are part of it. Millions of people who lend their computers to the search for extraterrestrial intelligence are part of it. The way buyers and sellers rate each other on Internet auction site eBay is part of it. Research by biologists, sociologists, and economists into the nature of cooperation offer explanatory frameworks. At least one key global business question is part of it - why is the Japanese company DoCoMo profiting from enhanced wireless Internet services while US and European mobile telephony operators struggle to avoid failure?&lt;br /&gt;
&lt;br /&gt;
The people who make up smart mobs cooperate in ways never before possible because they carry devices that possess both communication and computing capabilities. Their mobile devices connect them with other information devices in the environment as well as with other people&#039;s telephones. Dirt-cheap microprocessors embedded in everything from box tops to shoes are beginning to permeate furniture, buildings, neighborhoods, products with invisible intercommunicating smartifacts. When they connect the tangible objects and places of our daily lives with the Internet, handheld communication media mutate into wearable remote control devices for the physical world.&lt;br /&gt;
&lt;br /&gt;
Media cartels and government agencies are seeking to reimpose the regime of the broadcast era in which the customers of technology will be deprived of the power to create and left only with the power to consume. That power struggle is what the battles over file-sharing, copy-protection, regulation of the radio spectrum are about. Are the populations of tomorrow going to be users, like the PC owners and website creators who turned technology to widespread innovation? Or will they be consumers, constrained from innovation and locked into the technology and business models of the most powerful entrenched interests?&lt;br /&gt;
&lt;br /&gt;
Howard Rheingold is one of the world&#039;s foremost authorities on the social implications of technology. Over the past twenty years he has traveled around the world, observing and writing about emerging trends in computing, communications, and culture. One of the creators and former founding executive editor of HotWired, he has served as editor of The Whole Earth Review, editor-in-chief of The Millennium Whole Earth Catalog, and on-line host for The Well. The author of several books, including The Virtual Community, Virtual Reality, and Tools for Thought, he lives in Mill Valley, California.&lt;br /&gt;
&lt;br /&gt;
== Participating Meetup groups (5093 members) ==&lt;br /&gt;
&lt;br /&gt;
[[Atlanta Semantic Web Meetup|Atlanta]]  -  [[Austin Semantic Web Meetup|Austin]]  - [[Lotico CSW Berlin|Berlin]] -  [[Cambridge Semantic Web Meetup|Cambridge]]  -  [[The Chicago Semantic Web Meetup Group|Chicago]] - [[Frankfurt Semantic Web Meetup|Frankfurt]]  - [[Central Florida Semantic Web Meetup|Central Florida]] -  [[London Semantic Web Meetup|London]]  -  [[Los Angeles Semantic Web Meetup|Los Angeles]]  - [[Munchen-Semantic-Web-Meetup | Munich]] - [[SWNYC|New York]]  -  [[Philadelphia Semantic Web Meetup|Philadelphia]]  - [[Oslo Semantic Web Meetup|Oslo]]  - [[Ottawa Semantic Web Meetup|Ottawa]] -  [[Princeton Semantic Web Meetup|Princeton]]  -  [[San Diego Semantic Web Meetup|San Diego]]  -  [[San Francisco Semantic Web Meetup|San Francisco]]  - [[Santiago Semantic Web Meetup|Santiago]] - [[Seattle Semantic Web Meetup|Seattle]]  -  [[Silicon Valley Meetup|Silicon Valley]]  -  [[Thessaloniki Semantic Web Meetup|Thessaloniki]]  -  [[Toronto Semantic Web Meetup|Toronto]]  -  [[Vancouver Semantic Web Meetup|Vancouver]]  -  [[Vienna Semantic Web Meetup|Vienna]]  -  [[Washington Semantic Web Meetup|Washington DC]]&lt;br /&gt;
&lt;br /&gt;
== Proposed Sessions ==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TABLE CELLSPACING=3 CELLPADDING=3&amp;gt;&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Monday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 21&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#EEEEEE&amp;gt;Pre-conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt; &amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Tuesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 22&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;First Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Global Semantic Web Meetup&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Wednesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 23&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Second Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[[Semantic Code Camp 2010|Semantic Unconference]]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Thursday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 24&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Third Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Semantic Social Networks - Meetup Ontology 2pm-3pm Incentives &amp;amp; Roadblocks for Participating in the Semantic Web 7pm&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Friday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 25&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Fourth Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[http://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/calendar/13126881/ LODE: Linking Open Descriptions of Events]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/TABLE&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Proposed Session Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Incentives &amp;amp; Roadblocks for Participating in the Semantic Web]]&#039;&#039;&#039; - Griffin Caprio&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Semantic Code Camp 2010]]&#039;&#039;&#039; - Shamod Lacoul&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Big Meetup Social 2010|Global Meetup Social 2010]]&#039;&#039;&#039; - Marco Neumann&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6631</id>
		<title>Meetup at the Semantic Technology Conference 2010 in San Francisco</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_at_the_Semantic_Technology_Conference_2010_in_San_Francisco&amp;diff=6631"/>
		<updated>2026-04-04T08:44:40Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;__NOTOC__&lt;br /&gt;
&lt;br /&gt;
This year our global Lotico Semantic Web event took place at the Semantic Technology Conference 2010 in San Francisco. Come and meet your peers to discuss all things Semantic Web, Web 3.0 and Linked Data to make the Web of Data a reality. It took more than 10 years to get the Semantic Web initiative where it is today and we have good reason to believe that it&#039;s about time to hit the mainstream web. This I believe will not happen without friction since the standards in the Semantic Web initiative are geared towards a more academic audience rather than web practitioners. With the growing adoption in the mainstream community this might require some minor review of some recommendations, additional training documentation and more tools for development. So it&#039;s interesting times again, I hope to see you in San Francisco this summer.&lt;br /&gt;
&lt;br /&gt;
==Special Guest: Howard Rheingold - Semantic Social Networks and Smart Mobs==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;youtube width=80%&amp;gt;https://youtu.be/vATVkMTv8q8&amp;lt;/youtube&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[http://vimeo.com/20316782 Vimeo Video]&lt;br /&gt;
&lt;br /&gt;
Smart mobs emerge when communication and computing technologies amplify human talents for cooperation. The impacts of smart mob technology already appear to be both beneficial and destructive, used by some of its earliest adopters to support democracy and by others to coordinate terrorist attacks. The technologies that are beginning to make smart mobs possible are mobile communication devices and pervasive computing - inexpensive microprocessors embedded in everyday objects and environments. Already, governments have fallen, youth subcultures have blossomed from Asia to Scandinavia, new industries have been born and older industries have launched furious counterattacks.&lt;br /&gt;
&lt;br /&gt;
Street demonstrators in the 1999 anti-WTO protests used dynamically updated websites, cell-phones, and &amp;quot;swarming&amp;quot; tactics in the &amp;quot;battle of Seattle.&amp;quot; A million Filipinos toppled President Estrada through public demonstrations organized through salvos of text messages.&lt;br /&gt;
&lt;br /&gt;
The pieces of the puzzle are all around us now, but haven&#039;t joined together yet. The radio chips designed to replace barcodes on manufactured objects are part of it. Wireless Internet nodes in cafes, hotels, and neighborhoods are part of it. Millions of people who lend their computers to the search for extraterrestrial intelligence are part of it. The way buyers and sellers rate each other on Internet auction site eBay is part of it. Research by biologists, sociologists, and economists into the nature of cooperation offer explanatory frameworks. At least one key global business question is part of it - why is the Japanese company DoCoMo profiting from enhanced wireless Internet services while US and European mobile telephony operators struggle to avoid failure?&lt;br /&gt;
&lt;br /&gt;
The people who make up smart mobs cooperate in ways never before possible because they carry devices that possess both communication and computing capabilities. Their mobile devices connect them with other information devices in the environment as well as with other people&#039;s telephones. Dirt-cheap microprocessors embedded in everything from box tops to shoes are beginning to permeate furniture, buildings, neighborhoods, products with invisible intercommunicating smartifacts. When they connect the tangible objects and places of our daily lives with the Internet, handheld communication media mutate into wearable remote control devices for the physical world.&lt;br /&gt;
&lt;br /&gt;
Media cartels and government agencies are seeking to reimpose the regime of the broadcast era in which the customers of technology will be deprived of the power to create and left only with the power to consume. That power struggle is what the battles over file-sharing, copy-protection, regulation of the radio spectrum are about. Are the populations of tomorrow going to be users, like the PC owners and website creators who turned technology to widespread innovation? Or will they be consumers, constrained from innovation and locked into the technology and business models of the most powerful entrenched interests?&lt;br /&gt;
&lt;br /&gt;
Howard Rheingold is one of the world&#039;s foremost authorities on the social implications of technology. Over the past twenty years he has traveled around the world, observing and writing about emerging trends in computing, communications, and culture. One of the creators and former founding executive editor of HotWired, he has served as editor of The Whole Earth Review, editor-in-chief of The Millennium Whole Earth Catalog, and on-line host for The Well. The author of several books, including The Virtual Community, Virtual Reality, and Tools for Thought, he lives in Mill Valley, California.&lt;br /&gt;
&lt;br /&gt;
== Participating Meetup groups (5093 members) ==&lt;br /&gt;
&lt;br /&gt;
[[Atlanta Semantic Web Meetup|Atlanta]]  -  [[Austin Semantic Web Meetup|Austin]]  - [[Lotico CSW Berlin|Berlin]] -  [[Cambridge Semantic Web Meetup|Cambridge]]  -  [[The Chicago Semantic Web Meetup Group|Chicago]] - [[Frankfurt Semantic Web Meetup|Frankfurt]]  - [[Central Florida Semantic Web Meetup|Central Florida]] -  [[London Semantic Web Meetup|London]]  -  [[Los Angeles Semantic Web Meetup|Los Angeles]]  - [[Munchen-Semantic-Web-Meetup | Munich]] - [[SWNYC|New York]]  -  [[Philadelphia Semantic Web Meetup|Philadelphia]]  - [[Oslo Semantic Web Meetup|Oslo]]  - [[Ottawa Semantic Web Meetup|Ottawa]] -  [[Princeton Semantic Web Meetup|Princeton]]  -  [[San Diego Semantic Web Meetup|San Diego]]  -  [[San Francisco Semantic Web Meetup|San Francisco]]  - [[Santiago Semantic Web Meetup|Santiago]] - [[Seattle Semantic Web Meetup|Seattle]]  -  [[Silicon Valley Meetup|Silicon Valley]]  -  [[Thessaloniki Semantic Web Meetup|Thessaloniki]]  -  [[Toronto Semantic Web Meetup|Toronto]]  -  [[Vancouver Semantic Web Meetup|Vancouver]]  -  [[Vienna Semantic Web Meetup|Vienna]]  -  [[Washington Semantic Web Meetup|Washington DC]]&lt;br /&gt;
&lt;br /&gt;
== Proposed Sessions ==&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TABLE CELLSPACING=3 CELLPADDING=3&amp;gt;&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Monday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 21&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#EEEEEE&amp;gt;Pre-conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt; &amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Tuesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 22&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;First Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Global Semantic Web Meetup&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Wednesday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 23&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Second Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[[Semantic Code Camp 2010|Semantic Unconference]]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Thursday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 24&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Third Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;Semantic Social Networks - Meetup Ontology 2pm-3pm Incentives &amp;amp; Roadblocks for Participating in the Semantic Web 7pm&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;TR&amp;gt;&amp;lt;TD&amp;gt;Friday&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;June 25&amp;lt;/TD&amp;gt;&amp;lt;TD bgcolor=#DDDDDD&amp;gt;Fourth Conference Day&amp;lt;/TD&amp;gt;&amp;lt;TD&amp;gt;[http://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/calendar/13126881/ LODE: Linking Open Descriptions of Events]&amp;lt;/TD&amp;gt;&amp;lt;/TR&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/TABLE&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Proposed Session Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Incentives &amp;amp; Roadblocks for Participating in the Semantic Web]]&#039;&#039;&#039; - Griffin Caprio&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Semantic Code Camp 2010]]&#039;&#039;&#039; - Shamod Lacoul&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;[[Big Meetup Social 2010|Global Meetup Social 2010]]&#039;&#039;&#039; - Marco Neumann&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6630</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6630"/>
		<updated>2026-03-23T22:02:07Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;&amp;lt;youtube&amp;gt;https://youtu.be/&amp;lt;/youtube&amp;gt;&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
[https://en.wikipedia.org/wiki/Peter_Fox_(professor) http://www.rpi.edu/dept/ees/people/faculty/fox.html]&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Gale A. Brewer is the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
===Resources===&lt;br /&gt;
[[Data Gov Transcript]]&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6629</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6629"/>
		<updated>2026-03-23T21:58:21Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;&amp;lt;youtube&amp;gt;https://youtu.be/&amp;lt;/youtube&amp;gt;&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Gale A. Brewer is the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
===Resources===&lt;br /&gt;
[[Data Gov Transcript]]&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6628</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6628"/>
		<updated>2026-03-23T21:57:57Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;&amp;lt;youtube&amp;gt;https://youtu.be/&amp;lt;/youtube&amp;gt;&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Gale A. Brewer is the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Data Gov Transcript]]&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6627</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6627"/>
		<updated>2026-03-23T21:57:13Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;&amp;lt;youtube&amp;gt;https://youtu.be/&amp;lt;/youtube&amp;gt;&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Gale A. Brewer is the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6626</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6626"/>
		<updated>2026-03-23T21:56:42Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;video&amp;gt;&amp;lt;/video&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Gale A. Brewer is the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6625</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6625"/>
		<updated>2026-03-23T21:56:08Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Review */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Gale A. Brewer is the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6624</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6624"/>
		<updated>2026-03-23T15:08:36Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]] Lotico Director&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6623</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6623"/>
		<updated>2026-03-23T14:50:23Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Impressions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]]&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6622</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6622"/>
		<updated>2026-03-23T14:49:17Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Impressions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]]&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6621</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6621"/>
		<updated>2026-03-23T14:48:07Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]]&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&amp;lt;br&amp;gt;&amp;lt;/P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6620</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6620"/>
		<updated>2026-03-23T14:47:52Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]]&lt;br /&gt;
&lt;br /&gt;
&amp;lt;P&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6619</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6619"/>
		<updated>2026-03-23T14:47:40Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
Organizer: [[Marco Neumann]]&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6618</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6618"/>
		<updated>2026-03-23T14:44:47Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Review */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Conclusion:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6617</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6617"/>
		<updated>2026-03-23T14:44:25Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Review= */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review===&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
Conclusion: The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6616</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6616"/>
		<updated>2026-03-23T14:44:16Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
===Review====&lt;br /&gt;
&lt;br /&gt;
This meeting served as a critical intersection between municipal policy and advanced data science. At this time, New York City was on the precipice of a digital revolution. The event brought together Gale A. Brewer, the primary political architect of NYC’s transparency laws, and leading scientists from Rensselaer Polytechnic Institute (RPI) to discuss how &amp;quot;Linked Open Data&amp;quot; (LOD) could transform raw government records into actionable intelligence.&lt;br /&gt;
&lt;br /&gt;
Keynote: Open Data in Government&lt;br /&gt;
&lt;br /&gt;
Speaker: Gale A. Brewer, NYC Council Member &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Policy Objectives:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
In 2010, Brewer was the Chair of the Committee on Governmental Operations. Her presentation focused on the transition of government data from &amp;quot;locked&amp;quot; formats (PDFs and paper) to &amp;quot;machine-readable&amp;quot; formats. Key themes from her address included: The &amp;quot;Open Data Law&amp;quot; Vision: Brewer discussed the early stages of what would become Local Law 11 of 2012 (The NYC Open Data Law). She argued that government data belongs to the public and should be accessible without Freedom of Information Law (FOIL) requests. Efficiency and Accountability: She emphasized that open data isn&#039;t just for software developers; it is a tool for the Council to monitor agency performance, from pothole repairs to budget expenditures. Breaking Silos: A major point of her talk was the difficulty of inter-agency data sharing. She advocated for a centralized portal (which later became the NYC Open Data Portal) to standardize how different departments (NYPD, DOT, DOB) label and share information. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Scientific Deep Dive:&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
The Semantic Web FrameworkFollowing the policy introduction, the RPI delegation explained the technical solution to Brewer’s policy challenges: The Semantic Web. Technical Highlights: Jim Hendler (Data.gov): Hendler discussed his role in the federal Data.gov project. He explained that simply putting a CSV file online isn&#039;t enough; the data needs &amp;quot;context.&amp;quot; He introduced the concept of RDF (Resource Description Framework), which allows computers to understand the relationship between data points (e.g., recognizing that &amp;quot;Upper West Side&amp;quot; is a &amp;quot;Neighborhood&amp;quot; within &amp;quot;Manhattan&amp;quot;). &lt;br /&gt;
&lt;br /&gt;
Deborah McGuinness (Linked Data): She demonstrated how the LOD (Linked Open Data) Cloud allows different datasets to talk to each other. For example, linking NYC health department data with federal census data to identify trends in public health. Peter Fox (eScience): Fox expanded the scope to scientific collaboration, showing how the same semantic tools used for government transparency could be used by scientists to share massive climate and solar datasets across borders.4. Synthesis: Where Policy Meets ScienceThe transcript and event description highlight a specific synergy: Brewer provided the &amp;quot;What&amp;quot; (the data), and the RPI team provided the &amp;quot;How&amp;quot; (the Semantic Web). Challenge (Brewer) Solution (Hendler/McGuinness/Fox)Data in Silos: Agencies use different terms for the same thing. Ontologies: Creating a shared &amp;quot;vocabulary&amp;quot; so all systems recognize the same entities. Accessibility: Data is hard to find and use. URIs: Giving every piece of data a unique web address so it can be linked directly.Sustainability: High cost of manual data entry.Automation: Using machine-readable formats to allow software to update data in real-time. &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Historical Significance&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
This meeting is now viewed as a landmark moment for the NYC Open Data movement. Legislative Impact: Less than two years after this talk, Gale Brewer’s efforts culminated in the passage of the Open Data Law, making NYC the first city in the world to mandate that all public data be made available in a single web portal. Technological Legacy: The event solidified the role of the Semantic Web in public policy, moving the conversation from &amp;quot;Should we share data?&amp;quot; to &amp;quot;How do we make data intelligent?&amp;quot; &lt;br /&gt;
&lt;br /&gt;
Conclusion: The event was a successful bridge between the legislative world of City Hall and the high-tech world of semantic data, setting the stage for the modern &amp;quot;Smart City&amp;quot; infrastructure NYC uses today.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Gale_Brewer&amp;diff=6615</id>
		<title>Gale Brewer</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Gale_Brewer&amp;diff=6615"/>
		<updated>2026-03-23T08:48:50Z</updated>

		<summary type="html">&lt;p&gt;Marco: Created page with &amp;quot;  Category:Person&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&lt;br /&gt;
[[Category:Person]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6614</id>
		<title>Data Gov - Bringing Government and Scientific Data to the Web</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Data_Gov_-_Bringing_Government_and_Scientific_Data_to_the_Web&amp;diff=6614"/>
		<updated>2026-03-23T08:48:32Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Sep 2, 2010 · 6:00 PM&lt;br /&gt;
&lt;br /&gt;
Chapter: New York City&lt;br /&gt;
&lt;br /&gt;
Location: JPMorgan Chase&lt;br /&gt;
&lt;br /&gt;
Event ID: 14109835&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/semweb-25/events/14109835/&lt;br /&gt;
&lt;br /&gt;
Attendees: 111&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
At every level of government, electronic data has become increasingly central to those charged with forecasting and allocating funds for infrastructure and community projects. Now more than ever we depend on the availability and accessibility of data. And since this data is growing exponentially in size, we are confronted with critical questions of quality, reuseability and portability.&lt;br /&gt;
&lt;br /&gt;
In this event of the New York Semantic Web group we will explore datasets and applications which are build based on the semantic web framework to address these issues. Currently, a number of teams around the world translate open datasets into RDF to link local data to the so called Linked Open Data (LOD) cloud. Consequently the projects make this linked data the resource for the development of new applications and demos that consume linked open data.&lt;br /&gt;
&lt;br /&gt;
Data.gov is the primary source for much of the data that powers these applications to help smart people make better decisions in public and private organizations. This data pool is enriched with a number of publicly available datasets from industry, demographics and data from other countries and non-governmental organizations which will allow planners to make use of an unprecedented amount of reusable structured data for the decision making process.&lt;br /&gt;
&lt;br /&gt;
Schedule:&lt;br /&gt;
&lt;br /&gt;
6.00 pm Welcome&lt;br /&gt;
&lt;br /&gt;
6.15 pm Open Data in Government - [[Gale Brewer|Gale A. Brewer]], New York City Council Member&lt;br /&gt;
&lt;br /&gt;
6.30 pm [http://www.data.gov Data Gov] - Jim Hendler, Tetherless World Senior Constellation Professor, Rensselaer Polytechnic Institute (RPI)&lt;br /&gt;
&lt;br /&gt;
7.00 pm [http://files.meetup.com/274991/McGuinness_NYC_SWMeetup20100903Presented.pdf Semantic Web and Linked Data: Emerging Trends]. Deborah L. McGuinness, Tetherless World Senior Constellation Professor (RPI)&lt;br /&gt;
&lt;br /&gt;
7.30 pm [http://files.meetup.com/274991/Fox_NYC_SWMeetup20100902.pdf The eScience revolution: Semantic Web platforms for massive scientific collaboration] - Peter Fox, Professor and Tetherless World Research Constellation Chair (RPI)&lt;br /&gt;
&lt;br /&gt;
8.00 pm QA Panel&lt;br /&gt;
&lt;br /&gt;
8.30 pm Closing&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute to this session with your own story about publishing government data on the web with Semantic Web technologies please let us know.&lt;br /&gt;
&lt;br /&gt;
We would like to thank JPMorgan Chase for hosting the New York Semantic Web group. And for supporting the local efforts of Lotico, your global semantic social network.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;hr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Gale A. Brewer&lt;br /&gt;
&lt;br /&gt;
Council Member Gale A. Brewer has been representing the Upper West Side and the northern part of Clinton in the New York City Council since 2002. She was re-elected in the November 2009 general election with over 80 percent of the vote, receiving nearly 5,000 more votes than any of the other 50 Council Members who ran. In the November 2003 and 2005 elections, she received 86% and 80% of the vote, respectively. Her service in the Council is a continuation of nearly 30 years of public service. Brewer currently chairs the Committee on Governmental Operations. The Committee, consistent with its mandate, will review governmental structure and organization with an eye toward increasing both efficiency and accountability, particularly in the delivery of services and the use of technology. Brewer chaired the Committee on Technology in Government (now the Committee on Technology) from 2002-2009, where she worked to make better use of technology to save money, improve City services, and bring residents, businesses and non-profits closer to government and their communities.&lt;br /&gt;
Council Member Brewer looks forward to integrating many of the practices of the Technology in Government to the Governmental Operations Committee.&lt;br /&gt;
&lt;br /&gt;
http://council.nyc.gov/d6/html/members/home.shtml&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof James Hendler&lt;br /&gt;
&lt;br /&gt;
Jim Hendler is the Tetherless World Professor of Computer and Cognitive Science, and the Assistant Dean for Information Technology, at RPI. He is also an faculty affiliate of the Experimental Multimedia&lt;br /&gt;
Performing Arts Center (EMPAC) and he also serves as a Director of the international Web Science Research Initiative and is a visiting Professor at the Institute of Creative Technology at DeMontfort&lt;br /&gt;
University in Leicester, UK. Hendler has authored about 200 technical papers in the areas of artificial intelligence, Semantic Web, agent-based computing and high performance processing. One of the inventors of Semantic Web Hendler was the recipient of a 1995 Fulbright Foundation Fellowship, is a member of the US Air Force Science Advisory Board, and is a Fellow of the American Association for Artificial Intelligence and the British Computer Society. He is also the former Chief Scientist of the Information Systems Office at the US Defense Advanced Research Projects Agency (DARPA) and was awarded a US Air Force Exceptional Civilian Service Medal in 2002. He is the Editor-in-Chief emeritus of IEEE Intelligent Systems and is the first computer scientist to serve on the Board of Reviewing Editors&lt;br /&gt;
for Science.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~hendler/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Deborah McGuinness&lt;br /&gt;
&lt;br /&gt;
Deborah is the Tetherless World Senior Constellation Professor at the Department of Computer Science and Cognitive Science Department Rensselaer Polytechnic Institute (RPI). Before her new role at RPI Deborah was the acting director and senior research scientist at the Knowledge Systems, (KSL) Artificial Intelligence Laboratory at Stanford University. She is a leading expert in knowledge representation and reasoning languages and systems and has worked in ontology creation and evolution environments for over 20 years. Most recently, Deborah is best known for her leadership role in semantic web research, and for her work on explanation, trust, and applications of semantic web technology, particularly for scientific applications. Deborah was co-editor of the Ontology Web Language which has emerged from web ontology working group of the World Wide Web (W3C) semantic web activity and has now achieved W3C Recommendation status. She helped start the web ontology working group out of work as a co-author of the DARPA Agent Markup Language program&#039;s DAML language. She helped form the Joint EU/US Agent Markup Language Committee which evolved the DAML language into the DAML+OIL description logic-based ontology language. She is a co-author of one of the more widely used long-lived description logic systems (CLASSIC) from Bell Laboratories. Her work on languages (including OWL, DAML+OIL, OIL, CLASSIC, etc.) is aimed at providing languages that enable the next generation of web applications moving from a web aimed at human consumption to the semantic web aimed at machine consumption in support of intelligent&lt;br /&gt;
assistants and web agents. Deborah is a leader in ontology-based tools and applications.&lt;br /&gt;
&lt;br /&gt;
http://www.cs.rpi.edu/~dlm/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Prof Peter Fox&lt;br /&gt;
&lt;br /&gt;
Peter is the Tetherless World Research Constellation Chair for Climate Variability and Solar-Terrestrial Physics at RPI. He joined the Tetherless World Constellation in 2008. Formerly, he was the Chief Computational Scientist at the High Altitude Observatory (HAO) of the National Center for Atmospheric Research (NCAR). Fox&#039;s research specializes in the fields of solar and solar-terrestrial physics, computational and computer science, information technology, and grid-enabled, distributed semantic data frameworks. This research utilizes state-of-the-art modeling techniques, internet-based technologies, including the semantic web, and applies them to large-scale distributed scientific repositories addressing the full life-cycle of data and information within specific science and engineering disciplines as well as among disciplines.&lt;br /&gt;
&lt;br /&gt;
http://www.rpi.edu/dept/ees/people/faculty/fox.html&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/1.jpg http://www.lotico.com/images/20100902/thumb_1.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/2.jpg http://www.lotico.com/images/20100902/thumb_2.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/3.jpg http://www.lotico.com/images/20100902/thumb_3.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/4.jpg http://www.lotico.com/images/20100902/thumb_4.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/5.jpg http://www.lotico.com/images/20100902/thumb_5.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/6.jpg http://www.lotico.com/images/20100902/thumb_6.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/7.jpg http://www.lotico.com/images/20100902/thumb_7.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/8.jpg http://www.lotico.com/images/20100902/thumb_8.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/9.jpg http://www.lotico.com/images/20100902/thumb_9.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/10.jpg http://www.lotico.com/images/20100902/thumb_10.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/11.jpg http://www.lotico.com/images/20100902/thumb_11.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/12.jpg http://www.lotico.com/images/20100902/thumb_12.jpg]&lt;br /&gt;
&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/14.jpg http://www.lotico.com/images/20100902/thumb_14.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/15.jpg http://www.lotico.com/images/20100902/thumb_15.jpg]&lt;br /&gt;
[http://www.swnyc.org/images/img.php?img=http://www.lotico.com/images/20100902/16.jpg http://www.lotico.com/images/20100902/thumb_16.jpg]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=ISWC_Meetup_with_Neo4J&amp;diff=6613</id>
		<title>ISWC Meetup with Neo4J</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=ISWC_Meetup_with_Neo4J&amp;diff=6613"/>
		<updated>2026-02-16T15:54:57Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Video: [http://blip.tv/file/2895759 Lotico Meetup with Neo4J at ISWC 2009] &lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
[http://blip.tv/rss/flash/2915529 https://www.lotico.com/images/2009/Rashaun_Stovall_and_Marco.jpg]&lt;br /&gt;
&lt;br /&gt;
Date: October 27, 2009&lt;br /&gt;
&lt;br /&gt;
Video Edit: Ra&#039;Shaun Stovall&lt;br /&gt;
&lt;br /&gt;
Speaker: Emil Eifrem Neo4j&lt;br /&gt;
&lt;br /&gt;
Host: Marco Neumann KONA&lt;br /&gt;
&lt;br /&gt;
Location: [http://iswc2009.semanticweb.org/ ISWC 2009] - 14750 Conference Center Drive, Chantilly, Virginia 20151 USA&lt;br /&gt;
&lt;br /&gt;
==  Neo4j - The Benefits of Graph Databases and Property Graphs==&lt;br /&gt;
&lt;br /&gt;
Graph Database Systems &amp;amp; Neo4j - Exploring the relationship between Property Graph Systems and Semantic Web Databases. &lt;br /&gt;
&lt;br /&gt;
[[Emil Eifrém]], CEO Neo4j&lt;br /&gt;
&lt;br /&gt;
Description&lt;br /&gt;
&lt;br /&gt;
Meet [[Emil Eifrém]] CEO Neo Technology, founder of the Neo4j graph database project and CEO of Neo Technology. Programmer by passion the first 15 years on this planet and by passion &amp;amp; profession the remaining 15. First free software project at age 16. Now mainly focused on spreading the word about the powers of graphs and preaching the demise of tabular solutions everywhere. Presents regularly at conferences such as JAOO, Oredev, QCon, and OSCON.&lt;br /&gt;
&lt;br /&gt;
Many applications today handle data that is deeply associative, i.e. structured as graphs (networks). The most obvious example of this is social networking sites, but even tagging systems, content management systems and wikis deal with inherently hierarchical or graph-shaped data.&lt;br /&gt;
&lt;br /&gt;
This turns out to be a problem because it’s difficult to deal with recursive data structures in both traditional relational databases and many NoSQL stores. For example, in an RDBMS each traversal along a link in a graph is a join, and joins are known to be very expensive.&lt;br /&gt;
&lt;br /&gt;
A graph database uses nodes, relationships between nodes and key-value properties instead of tables to represent information. This model is typically substantially faster for associative data sets and uses a schema-less, bottoms-up model that is ideal for capturing ad-hoc and rapidly changing data.&lt;br /&gt;
&lt;br /&gt;
This session will introduce an open source, high-performance, transactional and disk-based graph database called “Neo4j”, which frequently outperforms relational backends with &amp;gt;1000x for many increasingly important use cases.&lt;br /&gt;
&lt;br /&gt;
===External Links===&lt;br /&gt;
http://neo4j.org&lt;br /&gt;
&lt;br /&gt;
[[category:event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=ISWC_Meetup_with_Neo4J&amp;diff=6612</id>
		<title>ISWC Meetup with Neo4J</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=ISWC_Meetup_with_Neo4J&amp;diff=6612"/>
		<updated>2026-02-16T15:54:40Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Video: [http://blip.tv/file/2895759 Lotico Meetup with Neo4J at ISWC 2009] [http://blip.tv/rss/flash/2915529 https://www.lotico.com/images/2009/Rashaun_Stovall_and_Marco.jpg]&lt;br /&gt;
&lt;br /&gt;
Date: October 27, 2009&lt;br /&gt;
&lt;br /&gt;
Video Edit: Ra&#039;Shaun Stovall&lt;br /&gt;
&lt;br /&gt;
Speaker: Emil Eifrem Neo4j&lt;br /&gt;
&lt;br /&gt;
Host: Marco Neumann KONA&lt;br /&gt;
&lt;br /&gt;
Location: [http://iswc2009.semanticweb.org/ ISWC 2009] - 14750 Conference Center Drive, Chantilly, Virginia 20151 USA&lt;br /&gt;
&lt;br /&gt;
==  Neo4j - The Benefits of Graph Databases and Property Graphs==&lt;br /&gt;
&lt;br /&gt;
Graph Database Systems &amp;amp; Neo4j - Exploring the relationship between Property Graph Systems and Semantic Web Databases. &lt;br /&gt;
&lt;br /&gt;
[[Emil Eifrém]], CEO Neo4j&lt;br /&gt;
&lt;br /&gt;
Description&lt;br /&gt;
&lt;br /&gt;
Meet [[Emil Eifrém]] CEO Neo Technology, founder of the Neo4j graph database project and CEO of Neo Technology. Programmer by passion the first 15 years on this planet and by passion &amp;amp; profession the remaining 15. First free software project at age 16. Now mainly focused on spreading the word about the powers of graphs and preaching the demise of tabular solutions everywhere. Presents regularly at conferences such as JAOO, Oredev, QCon, and OSCON.&lt;br /&gt;
&lt;br /&gt;
Many applications today handle data that is deeply associative, i.e. structured as graphs (networks). The most obvious example of this is social networking sites, but even tagging systems, content management systems and wikis deal with inherently hierarchical or graph-shaped data.&lt;br /&gt;
&lt;br /&gt;
This turns out to be a problem because it’s difficult to deal with recursive data structures in both traditional relational databases and many NoSQL stores. For example, in an RDBMS each traversal along a link in a graph is a join, and joins are known to be very expensive.&lt;br /&gt;
&lt;br /&gt;
A graph database uses nodes, relationships between nodes and key-value properties instead of tables to represent information. This model is typically substantially faster for associative data sets and uses a schema-less, bottoms-up model that is ideal for capturing ad-hoc and rapidly changing data.&lt;br /&gt;
&lt;br /&gt;
This session will introduce an open source, high-performance, transactional and disk-based graph database called “Neo4j”, which frequently outperforms relational backends with &amp;gt;1000x for many increasingly important use cases.&lt;br /&gt;
&lt;br /&gt;
===External Links===&lt;br /&gt;
http://neo4j.org&lt;br /&gt;
&lt;br /&gt;
[[category:event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_with_Sir_Tim_Berners-Lee_2009&amp;diff=6611</id>
		<title>Meetup with Sir Tim Berners-Lee 2009</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_with_Sir_Tim_Berners-Lee_2009&amp;diff=6611"/>
		<updated>2026-02-15T16:43:37Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Impressions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[http://www.w3c.org http://www.lotico.com/images/tim.jpg]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Chapter: San Francisco&lt;br /&gt;
&lt;br /&gt;
Date: October 27, 2009&lt;br /&gt;
&lt;br /&gt;
Event ID: 11017898&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/events/11017898/&lt;br /&gt;
&lt;br /&gt;
Meetup with [[Tim Berners-Lee]], Director [http://www.w3c.org W3C]  and [[Marco Neumann]] the organizer of Lotico the Semantic Web Meetup Alliance and local organizer in New York and San Francisco on Tuesday, October 27. The event will bring as many as possible organizers and meetup members together in one location in one night. So far we have the following organizers confirmed Barbara Starr (San Diego, CA), Christine Connors (Princeton, NJ), Juan Sequeda (Austin, TX), Lee Feigenbaum (Cambridge, MA), Markus Luczak-Rösch (Berlin, DE), Morton Swimmer (Zuerich, CH), Benjamin Grosof (Seattle, WA), Jamie Taylor (San Francisco,CA) and our local organizer in Washington DC, Brian Eubanks.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Location:&lt;br /&gt;
 Westfields Marriott&lt;br /&gt;
 14750 Conference Center Dr&lt;br /&gt;
 Chantilly, VA 20151&lt;br /&gt;
&lt;br /&gt;
==Video==&lt;br /&gt;
&lt;br /&gt;
http://vimeo.com/7459091&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[https://www.lotico.com/images/2009/10/2-300x199.jpg https://www.lotico.com/images/2009/10/2-300x199.jpg][https://www.lotico.com/images/2009/10/1-300x199.jpg https://www.lotico.com/images/2009/10/1-300x199.jpg][https://www.lotico.com/images/2009/11/4.jpg https://www.lotico.com/images/2009/11/4_300.jpg]&lt;br /&gt;
&lt;br /&gt;
[https://www.lotico.com/images/2009/11/6.JPG https://www.lotico.com/images/2009/11/6_300.JPG][https://www.lotico.com/images/2009/11/7.JPG https://www.lotico.com/images/2009/11/7_300.JPG][https://www.lotico.com/images/2009/11/8.JPG https://www.lotico.com/images/2009/11/8_300.JPG]&lt;br /&gt;
&lt;br /&gt;
[https://www.lotico.com/images/2009/11/9.JPG https://www.lotico.com/images/2009/11/9_300.JPG][https://www.lotico.com/images/2009/11/13.jpg https://www.lotico.com/images/2009/11/13_300.jpg][https://www.lotico.com/images/2009/11/11.JPG https://www.lotico.com/images/2009/11/11_300.JPG]&lt;br /&gt;
&lt;br /&gt;
[https://www.lotico.com/images/2009/11/tim_marco.jpg https://www.lotico.com/images/2009/11/12_300.JPG][https://www.lotico.com/images/2009/11/14.jpg https://www.lotico.com/images/2009/11/14_300.jpg][https://www.lotico.com/images/2009/11/19.jpg https://www.lotico.com/images/2009/11/19_300.jpg]&lt;br /&gt;
&lt;br /&gt;
[https://www.lotico.com/images/2009/11/KONA20091027_EGT.jpg https://www.lotico.com/images/2009/11/KONA20091027_EGT_300.jpg][https://www.lotico.com/images/2009/11/5.JPG https://www.lotico.com/images/2009/11/5_300.JPG][https://www.lotico.com/images/2009/11/16.jpg https://www.lotico.com/images/2009/11/16_300.jpg]&lt;br /&gt;
&lt;br /&gt;
[https://www.lotico.com/images/2009/11/17.jpg https://www.lotico.com/images/2009/11/17_300.jpg][https://www.lotico.com/images/2009/11/18.jpg https://www.lotico.com/images/2009/11/18_300.jpg][https://www.lotico.com/images/2009/11/10.JPG https://www.lotico.com/images/2009/11/10_300.JPG]&lt;br /&gt;
&lt;br /&gt;
[https://www.lotico.com/images/2009/11/20.jpg https://www.lotico.com/images/2009/11/20_300.jpg][https://www.lotico.com/images/2009/11/15.jpg https://www.lotico.com/images/2009/11/15_300.jpg]&lt;br /&gt;
[https://www.lotico.com/images/2009/11/3.jpg https://www.lotico.com/images/2009/11/3_300.jpg]&lt;br /&gt;
&lt;br /&gt;
==Blog Refs==&lt;br /&gt;
Semantic Social Networks – A Meetup with 400 people and Tim Berners-Lee&amp;lt;br&amp;gt;&lt;br /&gt;
http://www.marconeumann.org/blog/?p=23&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Meetup_with_Sir_Tim_Berners-Lee_2009&amp;diff=6610</id>
		<title>Meetup with Sir Tim Berners-Lee 2009</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Meetup_with_Sir_Tim_Berners-Lee_2009&amp;diff=6610"/>
		<updated>2026-02-15T16:42:59Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[http://www.w3c.org http://www.lotico.com/images/tim.jpg]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Chapter: San Francisco&lt;br /&gt;
&lt;br /&gt;
Date: October 27, 2009&lt;br /&gt;
&lt;br /&gt;
Event ID: 11017898&lt;br /&gt;
&lt;br /&gt;
URL: https://www.meetup.com/The-San-Francisco-Semantic-Web-Meetup/events/11017898/&lt;br /&gt;
&lt;br /&gt;
Meetup with [[Tim Berners-Lee]], Director [http://www.w3c.org W3C]  and [[Marco Neumann]] the organizer of Lotico the Semantic Web Meetup Alliance and local organizer in New York and San Francisco on Tuesday, October 27. The event will bring as many as possible organizers and meetup members together in one location in one night. So far we have the following organizers confirmed Barbara Starr (San Diego, CA), Christine Connors (Princeton, NJ), Juan Sequeda (Austin, TX), Lee Feigenbaum (Cambridge, MA), Markus Luczak-Rösch (Berlin, DE), Morton Swimmer (Zuerich, CH), Benjamin Grosof (Seattle, WA), Jamie Taylor (San Francisco,CA) and our local organizer in Washington DC, Brian Eubanks.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Location:&lt;br /&gt;
 Westfields Marriott&lt;br /&gt;
 14750 Conference Center Dr&lt;br /&gt;
 Chantilly, VA 20151&lt;br /&gt;
&lt;br /&gt;
==Video==&lt;br /&gt;
&lt;br /&gt;
http://vimeo.com/7459091&lt;br /&gt;
&lt;br /&gt;
==Impressions==&lt;br /&gt;
&lt;br /&gt;
[images/2009/10/2-300x199.jpg images/2009/10/2-300x199.jpg][images/2009/10/1-300x199.jpg images/2009/10/1-300x199.jpg][images/2009/11/4.jpg images/2009/11/4_300.jpg]&lt;br /&gt;
&lt;br /&gt;
[images/2009/11/6.JPG images/2009/11/6_300.JPG][images/2009/11/7.JPG images/2009/11/7_300.JPG][images/2009/11/8.JPG images/2009/11/8_300.JPG]&lt;br /&gt;
&lt;br /&gt;
[images/2009/11/9.JPG images/2009/11/9_300.JPG][images/2009/11/13.jpg images/2009/11/13_300.jpg][images/2009/11/11.JPG images/2009/11/11_300.JPG]&lt;br /&gt;
&lt;br /&gt;
[images/2009/11/tim_marco.jpg images/2009/11/12_300.JPG][images/2009/11/14.jpg images/2009/11/14_300.jpg][images/2009/11/19.jpg images/2009/11/19_300.jpg]&lt;br /&gt;
&lt;br /&gt;
[images/2009/11/KONA20091027_EGT.jpg images/2009/11/KONA20091027_EGT_300.jpg][images/2009/11/5.JPG images/2009/11/5_300.JPG][images/2009/11/16.jpg images/2009/11/16_300.jpg]&lt;br /&gt;
&lt;br /&gt;
[images/2009/11/17.jpg images/2009/11/17_300.jpg][images/2009/11/18.jpg images/2009/11/18_300.jpg][images/2009/11/10.JPG images/2009/11/10_300.JPG]&lt;br /&gt;
&lt;br /&gt;
[images/2009/11/20.jpg images/2009/11/20_300.jpg][images/2009/11/15.jpg images/2009/11/15_300.jpg]&lt;br /&gt;
[images/2009/11/3.jpg images/2009/11/3_300.jpg]&lt;br /&gt;
&lt;br /&gt;
==Blog Refs==&lt;br /&gt;
Semantic Social Networks – A Meetup with 400 people and Tim Berners-Lee&amp;lt;br&amp;gt;&lt;br /&gt;
http://www.marconeumann.org/blog/?p=23&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Semantics,_Blogs_and_Linked_Data_Technologies_In_The_Enterprise&amp;diff=6609</id>
		<title>Semantics, Blogs and Linked Data Technologies In The Enterprise</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Semantics,_Blogs_and_Linked_Data_Technologies_In_The_Enterprise&amp;diff=6609"/>
		<updated>2026-01-29T15:19:37Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Wednesday, August 8, 2012 - 6:30 PM to 9:30 PM PDT&lt;br /&gt;
&lt;br /&gt;
Location: Federated Media Publishing Inc, 72 Townsend Street · San Francisco, CA&lt;br /&gt;
&lt;br /&gt;
===Exploring Semantic Web Data - Roberto García===&lt;br /&gt;
http://rhizomik.net/~roberto/&lt;br /&gt;
&lt;br /&gt;
Roberto will present the Rhizomer tool, intended for semantic data exploration. This tool automatically generates a user interface tailored to a dataset and uses Information Architecture components users are already used to. These components are enriched with some additional functionalities that allow profiting from highly structured data and build complex semantic queries without requiring any knowledge about the underlying technologies or the terms used in the dataset. [[Roberto García]] is an associate professor at Universitat de Lleida (Spain) and currently visiting the Stanford HCI Group. see link for more info http://rhizomik.net/rhizomer/&lt;br /&gt;
&lt;br /&gt;
https://rhizomik.net/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;i&amp;gt;Full presentations&amp;lt;/i&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Semantics &amp;amp; Blogs - Tim Musgrove===&lt;br /&gt;
http://about.me/tmusgrove&lt;br /&gt;
&lt;br /&gt;
In this presentation Tim will present recent case studies that he has done, which are potentially of interest to everyone applying semantics to blogs. Such as: why the need for semantics in advertising is greater for blogs than it is for other types of content sites. And: how semantics analysis can help you understand why certain blog posts are tweeted/shared/liked more than others (with some surprising findings).&lt;br /&gt;
&lt;br /&gt;
[[Tim Musgrove]] is Chief Scientist at Federated Media Publishing, Inc (http://www.federatedmedia.net)&lt;br /&gt;
&lt;br /&gt;
Federated Media Publishing powers the Independent Web, and believes that the majority of meaningful engagements across digital media occur via high-quality independent sites and services. These sites leverage top digital talent to attract influential audiences who together create meaningful dialogue. Brands benefit from improved loyalty and increased sales when they become part of this authentic experience. Learn more at www.federatedmedia.net (http://www.federatedmedia.net/)&lt;br /&gt;
&lt;br /&gt;
===Linked data technologies meet enterprises - with PoolParty - Florian Kondert===&lt;br /&gt;
http://www.semantic-web.at/&lt;br /&gt;
&lt;br /&gt;
PoolParty Suite makes rigorous use of open standards and extensively uses linked data technologies. Thus, the way towards the web of data is open as well for enterprises, but not only to achieve better network effects within their ecosystems but also to create and maintain knowledge models more efficiently. http://poolparty.biz (http://poolparty.biz/)&lt;br /&gt;
&lt;br /&gt;
The presentation will include:&lt;br /&gt;
A brief overview about the architecture of PoolParty&lt;br /&gt;
Auto-population functionality for knowledge models&lt;br /&gt;
The integrated linked data interface and frontend&lt;br /&gt;
Linked data based synonyms and translation services&lt;br /&gt;
Customers examples&lt;br /&gt;
Q&amp;amp;A&lt;br /&gt;
&lt;br /&gt;
About PoolParty Suite&lt;br /&gt;
&lt;br /&gt;
PoolParty has outstanding capabilities to extract meaning from huge information sources based on W3C standard SKOS and linked data technologies. The software supports global 500 companies entering the semantic age of information management with scalable, easy to manage, corporate knowledge models.&lt;br /&gt;
&lt;br /&gt;
PoolParty APIs allow integration of semantic technologies with other systems like CMS, DMS, intranets or collaboration software like Confluence, Drupal and Sharepoint.&lt;br /&gt;
PoolParty is used by (to name a few) Pearson PLC, Roche Diagnostics, Credit Suisse, British Museum, Wolters Kluwer, Cornell University, The World Bank, Education Services Australia.&lt;br /&gt;
&lt;br /&gt;
About [[Florian Kondert]]&lt;br /&gt;
&lt;br /&gt;
Florian joined PoolParty Team early 2011 as COO and Head of Business Development. His educational and professional background in corporate communications, knowledge management, international sales and intercultural affairs enforces Florian&#039;s ability to transport complex issues to all professional levels from CXOs to end-users.&lt;br /&gt;
&lt;br /&gt;
Florian&#039;s main focus is to attend interested parties at a very early project-stage to design suitable approaches and business driven perspectives on semantic information management and hands-on use cases.&lt;br /&gt;
&lt;br /&gt;
(your registration fee will be used to confirm your registration and to order your pizza, snacks &amp;amp; drinks)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Semantics,_Blogs_and_Linked_Data_Technologies_In_The_Enterprise&amp;diff=6608</id>
		<title>Semantics, Blogs and Linked Data Technologies In The Enterprise</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Semantics,_Blogs_and_Linked_Data_Technologies_In_The_Enterprise&amp;diff=6608"/>
		<updated>2026-01-29T15:19:20Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Exploring Semantic Web Data - Roberto García */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Wednesday, August 8, 2012 - 6:30 PM to 9:30 PM PDT&lt;br /&gt;
&lt;br /&gt;
Location: Federated Media Publishing Inc, 72 Townsend Street · San Francisco, CA&lt;br /&gt;
&lt;br /&gt;
===Exploring Semantic Web Data - Roberto García===&lt;br /&gt;
http://rhizomik.net/~roberto/&lt;br /&gt;
&lt;br /&gt;
Roberto will present the Rhizomer tool, intended for semantic data exploration. This tool automatically generates a user interface tailored to the dataset and uses Information Architecture components users are already used to. These components are enriched with some additional functionalities that allow profiting from highly structured data and build complex semantic queries without requiring any knowledge about the underlying technologies or the terms used in the dataset. [[Roberto García]] is an associate professor at Universitat de Lleida (Spain) and currently visiting the Stanford HCI Group. see link for more info http://rhizomik.net/rhizomer/&lt;br /&gt;
&lt;br /&gt;
https://rhizomik.net/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;i&amp;gt;Full presentations&amp;lt;/i&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Semantics &amp;amp; Blogs - Tim Musgrove===&lt;br /&gt;
http://about.me/tmusgrove&lt;br /&gt;
&lt;br /&gt;
In this presentation Tim will present recent case studies that he has done, which are potentially of interest to everyone applying semantics to blogs. Such as: why the need for semantics in advertising is greater for blogs than it is for other types of content sites. And: how semantics analysis can help you understand why certain blog posts are tweeted/shared/liked more than others (with some surprising findings).&lt;br /&gt;
&lt;br /&gt;
[[Tim Musgrove]] is Chief Scientist at Federated Media Publishing, Inc (http://www.federatedmedia.net)&lt;br /&gt;
&lt;br /&gt;
Federated Media Publishing powers the Independent Web, and believes that the majority of meaningful engagements across digital media occur via high-quality independent sites and services. These sites leverage top digital talent to attract influential audiences who together create meaningful dialogue. Brands benefit from improved loyalty and increased sales when they become part of this authentic experience. Learn more at www.federatedmedia.net (http://www.federatedmedia.net/)&lt;br /&gt;
&lt;br /&gt;
===Linked data technologies meet enterprises - with PoolParty - Florian Kondert===&lt;br /&gt;
http://www.semantic-web.at/&lt;br /&gt;
&lt;br /&gt;
PoolParty Suite makes rigorous use of open standards and extensively uses linked data technologies. Thus, the way towards the web of data is open as well for enterprises, but not only to achieve better network effects within their ecosystems but also to create and maintain knowledge models more efficiently. http://poolparty.biz (http://poolparty.biz/)&lt;br /&gt;
&lt;br /&gt;
The presentation will include:&lt;br /&gt;
A brief overview about the architecture of PoolParty&lt;br /&gt;
Auto-population functionality for knowledge models&lt;br /&gt;
The integrated linked data interface and frontend&lt;br /&gt;
Linked data based synonyms and translation services&lt;br /&gt;
Customers examples&lt;br /&gt;
Q&amp;amp;A&lt;br /&gt;
&lt;br /&gt;
About PoolParty Suite&lt;br /&gt;
&lt;br /&gt;
PoolParty has outstanding capabilities to extract meaning from huge information sources based on W3C standard SKOS and linked data technologies. The software supports global 500 companies entering the semantic age of information management with scalable, easy to manage, corporate knowledge models.&lt;br /&gt;
&lt;br /&gt;
PoolParty APIs allow integration of semantic technologies with other systems like CMS, DMS, intranets or collaboration software like Confluence, Drupal and Sharepoint.&lt;br /&gt;
PoolParty is used by (to name a few) Pearson PLC, Roche Diagnostics, Credit Suisse, British Museum, Wolters Kluwer, Cornell University, The World Bank, Education Services Australia.&lt;br /&gt;
&lt;br /&gt;
About [[Florian Kondert]]&lt;br /&gt;
&lt;br /&gt;
Florian joined PoolParty Team early 2011 as COO and Head of Business Development. His educational and professional background in corporate communications, knowledge management, international sales and intercultural affairs enforces Florian&#039;s ability to transport complex issues to all professional levels from CXOs to end-users.&lt;br /&gt;
&lt;br /&gt;
Florian&#039;s main focus is to attend interested parties at a very early project-stage to design suitable approaches and business driven perspectives on semantic information management and hands-on use cases.&lt;br /&gt;
&lt;br /&gt;
(your registration fee will be used to confirm your registration and to order your pizza, snacks &amp;amp; drinks)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Semantics,_Blogs_and_Linked_Data_Technologies_In_The_Enterprise&amp;diff=6607</id>
		<title>Semantics, Blogs and Linked Data Technologies In The Enterprise</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Semantics,_Blogs_and_Linked_Data_Technologies_In_The_Enterprise&amp;diff=6607"/>
		<updated>2026-01-29T14:55:47Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Date: Wednesday, August 8, 2012 - 6:30 PM to 9:30 PM PDT&lt;br /&gt;
&lt;br /&gt;
Location: Federated Media Publishing Inc, 72 Townsend Street · San Francisco, CA&lt;br /&gt;
&lt;br /&gt;
===Exploring Semantic Web Data - Roberto García===&lt;br /&gt;
http://rhizomik.net/~roberto/&lt;br /&gt;
&lt;br /&gt;
Roberto will present the Rhizomer tool, intended for semantic data exploration. This tool automatically generates an user interface tailored to the dataset and uses Information Architecture components users are already used to. These components are enriched with some additional functionalities that allow profiting from highly structured data and build complex semantic queries without requiring any knowledge about the underlying technologies or the terms used in the dataset. [[Roberto García]] is an associate professor at Universitat de Lleida (Spain) and currently visiting the Stanford HCI Group. see link for more info http://rhizomik.net/rhizomer/&lt;br /&gt;
&lt;br /&gt;
https://rhizomik.net/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&amp;lt;i&amp;gt;Full presentations&amp;lt;/i&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Semantics &amp;amp; Blogs - Tim Musgrove===&lt;br /&gt;
http://about.me/tmusgrove&lt;br /&gt;
&lt;br /&gt;
In this presentation Tim will present recent case studies that he has done, which are potentially of interest to everyone applying semantics to blogs. Such as: why the need for semantics in advertising is greater for blogs than it is for other types of content sites. And: how semantics analysis can help you understand why certain blog posts are tweeted/shared/liked more than others (with some surprising findings).&lt;br /&gt;
&lt;br /&gt;
[[Tim Musgrove]] is Chief Scientist at Federated Media Publishing, Inc (http://www.federatedmedia.net)&lt;br /&gt;
&lt;br /&gt;
Federated Media Publishing powers the Independent Web, and believes that the majority of meaningful engagements across digital media occur via high-quality independent sites and services. These sites leverage top digital talent to attract influential audiences who together create meaningful dialogue. Brands benefit from improved loyalty and increased sales when they become part of this authentic experience. Learn more at www.federatedmedia.net (http://www.federatedmedia.net/)&lt;br /&gt;
&lt;br /&gt;
===Linked data technologies meet enterprises - with PoolParty - Florian Kondert===&lt;br /&gt;
http://www.semantic-web.at/&lt;br /&gt;
&lt;br /&gt;
PoolParty Suite makes rigorous use of open standards and extensively uses linked data technologies. Thus, the way towards the web of data is open as well for enterprises, but not only to achieve better network effects within their ecosystems but also to create and maintain knowledge models more efficiently. http://poolparty.biz (http://poolparty.biz/)&lt;br /&gt;
&lt;br /&gt;
The presentation will include:&lt;br /&gt;
A brief overview about the architecture of PoolParty&lt;br /&gt;
Auto-population functionality for knowledge models&lt;br /&gt;
The integrated linked data interface and frontend&lt;br /&gt;
Linked data based synonyms and translation services&lt;br /&gt;
Customers examples&lt;br /&gt;
Q&amp;amp;A&lt;br /&gt;
&lt;br /&gt;
About PoolParty Suite&lt;br /&gt;
&lt;br /&gt;
PoolParty has outstanding capabilities to extract meaning from huge information sources based on W3C standard SKOS and linked data technologies. The software supports global 500 companies entering the semantic age of information management with scalable, easy to manage, corporate knowledge models.&lt;br /&gt;
&lt;br /&gt;
PoolParty APIs allow integration of semantic technologies with other systems like CMS, DMS, intranets or collaboration software like Confluence, Drupal and Sharepoint.&lt;br /&gt;
PoolParty is used by (to name a few) Pearson PLC, Roche Diagnostics, Credit Suisse, British Museum, Wolters Kluwer, Cornell University, The World Bank, Education Services Australia.&lt;br /&gt;
&lt;br /&gt;
About [[Florian Kondert]]&lt;br /&gt;
&lt;br /&gt;
Florian joined PoolParty Team early 2011 as COO and Head of Business Development. His educational and professional background in corporate communications, knowledge management, international sales and intercultural affairs enforces Florian&#039;s ability to transport complex issues to all professional levels from CXOs to end-users.&lt;br /&gt;
&lt;br /&gt;
Florian&#039;s main focus is to attend interested parties at a very early project-stage to design suitable approaches and business driven perspectives on semantic information management and hands-on use cases.&lt;br /&gt;
&lt;br /&gt;
(your registration fee will be used to confirm your registration and to order your pizza, snacks &amp;amp; drinks)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=E_for_Events_-_Harnessing_the_Disagreement_with_Lora_Aroyo_@_Tagasauris&amp;diff=6606</id>
		<title>E for Events - Harnessing the Disagreement with Lora Aroyo @ Tagasauris</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=E_for_Events_-_Harnessing_the_Disagreement_with_Lora_Aroyo_@_Tagasauris&amp;diff=6606"/>
		<updated>2026-01-25T18:54:53Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Tagasauris with Todd Carter */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Session-Level: Intermediate-Advanced&amp;lt;br&amp;gt;&lt;br /&gt;
Session-Type: Technology-NLP-New Ideas-Tagging-Media-Linked Data&lt;br /&gt;
&lt;br /&gt;
Date: October 18, 2012&amp;lt;br&amp;gt;&lt;br /&gt;
Location: Tagasauris, 175 Varick Street , 10014, New York, NY&lt;br /&gt;
&lt;br /&gt;
Meetup: [http://www.meetup.com/semweb-25/events/84799412/ meetup.com]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==E for Events - Harnessing the Disagreement with Lora Aroyo==&lt;br /&gt;
&lt;br /&gt;
At this session we will take a closer look at events-driven information extraction, data presentation, collection management and metadata enrichment with Lora Aroyo. An ever increasing amount of digital content is unleashed on the Web on a daily basis. This introduces a number of challenges for data managers and end users, e.g. the existing collection metadata process is not sufficient and often not suitable to support the range of user needs and interactions. Similarly collection vocabularies miss links to shared knowledge and the perspective of end users. In this context, we observe a shift from well-curated and closed environments, e.g. museum exhibitions, guide tours, where users were exposed only to a pre-selected set of objects, carefully arranged in stories and narratives. Nowadays, the users have access to an endless pool of information, but the meaning of the individual objects is lost, because of often incomplete descriptions, lack of links between the objects and different collections.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Lora Aroyo is an IBM faculty award winner 2012 for Computer Science. She is an Associate professor for Web and Media at the VU Amsterdam and Scientific coordinator of EU Projects such as NoTube, for integration of Web and TV data with the help of semantics and CHIP, for Cultural Heritage Information Personalization. She is currently on at IBM Research. &lt;br /&gt;
Her research focuses on using semantic web technologies for modeling user interests and context and applying them in recommendation systems and personalized access to online cultural heritage collections, multimedia archives and interactive TV. She has coordinated the CHIP project on Cultural Heritage Information Personalization and the NoTube project on the integration of Web &amp;amp; TV with the help of semantics. She has co-organized numerous workshops on personalized access to cultural heritage, e-learning, interactive television, visual interfaces to the social and semantic web (e.g. PATCH, FutureTV, PersWeb, VISSW and DeRIVE). Lora is actively involved in the Semantic Web community, i.e. program co-chair for ESWC2009 and ISWC2011 and conference chair for ESWC2010. She is also actively involved in the Personalization and User modeling community as vice-president of UM Inc. and a member of the editorial board for the UMUAI journal.&lt;br /&gt;
&lt;br /&gt;
For more info: http://lora-aroyo.org&amp;lt;br&amp;gt;&lt;br /&gt;
For slides: http://www.slideshare.net/laroyo&amp;lt;br&amp;gt;&lt;br /&gt;
Follow me on twitter: @laroyo&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Tagasauris with Todd Carter==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Tagasauris the smart way to tag. Tagasauris is a start up based in New York City that develops innovative services and products .&lt;br /&gt;
&lt;br /&gt;
[[Todd Carter]] is the CEO and Co-founder of Tagasauris, a meta-data curation platform that makes your content smarter! We combine crowdsourcing, machine learning and semantic technologies that increases the lifetime value of digital media by making it more discoverable, connected and engaging. Todd is widely respected as a creative visionary and leader in the digital asset, photography and linked open data community, with over 20 years experience working with photo archives, libraries, museums and information technology systems. Tagasauris has been featured in The New York Times, Wired, Business Week, The Economist and others. The National Endowment for the Humanities awarded Tagasauris and The Museum of the City of New York a grant to annotate the museum&#039;s archive. Tagasauris is launching it&#039;s first photo tagging app that promises to be a game-changer in the media and entertainment industry. Tagasauris was founded in December 2010 with headquarters in New York City.&lt;br /&gt;
&lt;br /&gt;
Dealing with raw data at scale presents present a grand challenge. This is particularly true for visual media where the the large disparity between descriptions of multimedia content that can be computed automatically and the richness and subjectivity of semantics used to find, organize and share visual media present what is called the “semantic gap”.&lt;br /&gt;
One way to the address this gap would be to relabel raw data. Use experts. Describe the features of visual media with information about their content. But this is impossibly slow and prohibitively expensive! There has to be a better way.&lt;br /&gt;
Tagasauris has built a large-scale human computation engine that addresses the grand challenge posed by the semantic gap. The tagasauris process automates the discovery and generation of semantically linked metadata and enables us to quickly and cost effectively build a new, interlinked data layer of context for multimedia on the web that &amp;quot;fills in&amp;quot; the semantic gap.&lt;br /&gt;
The power of this layer will come from the quantity and quality of links between media objects on the visual web and their intersection other objects like events that can be detected and modeled from within the interest and social graphs. &lt;br /&gt;
We’re interested in discussion and exploration of uses cases that exploit the power of a personalized and semantically interlinked visual web.&lt;br /&gt;
For more information please visit www.tagasauris.com or email: info@tagasauris.com&lt;br /&gt;
&lt;br /&gt;
https://tagasauris.com/&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=Todd_Carter&amp;diff=6605</id>
		<title>Todd Carter</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=Todd_Carter&amp;diff=6605"/>
		<updated>2026-01-25T18:50:58Z</updated>

		<summary type="html">&lt;p&gt;Marco: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Todd Carter]] is the CEO and Co-founder of Tagasauris, a meta-data curation platform that makes your content smarter!&lt;br /&gt;
&lt;br /&gt;
[[Category:Person]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
	<entry>
		<id>https://www.lotico.com/index.php?title=E_for_Events_-_Harnessing_the_Disagreement_with_Lora_Aroyo_@_Tagasauris&amp;diff=6604</id>
		<title>E for Events - Harnessing the Disagreement with Lora Aroyo @ Tagasauris</title>
		<link rel="alternate" type="text/html" href="https://www.lotico.com/index.php?title=E_for_Events_-_Harnessing_the_Disagreement_with_Lora_Aroyo_@_Tagasauris&amp;diff=6604"/>
		<updated>2026-01-25T18:50:49Z</updated>

		<summary type="html">&lt;p&gt;Marco: /* Tagasauris with Todd Carter */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Session-Level: Intermediate-Advanced&amp;lt;br&amp;gt;&lt;br /&gt;
Session-Type: Technology-NLP-New Ideas-Tagging-Media-Linked Data&lt;br /&gt;
&lt;br /&gt;
Date: October 18, 2012&amp;lt;br&amp;gt;&lt;br /&gt;
Location: Tagasauris, 175 Varick Street , 10014, New York, NY&lt;br /&gt;
&lt;br /&gt;
Meetup: [http://www.meetup.com/semweb-25/events/84799412/ meetup.com]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==E for Events - Harnessing the Disagreement with Lora Aroyo==&lt;br /&gt;
&lt;br /&gt;
At this session we will take a closer look at events-driven information extraction, data presentation, collection management and metadata enrichment with Lora Aroyo. An ever increasing amount of digital content is unleashed on the Web on a daily basis. This introduces a number of challenges for data managers and end users, e.g. the existing collection metadata process is not sufficient and often not suitable to support the range of user needs and interactions. Similarly collection vocabularies miss links to shared knowledge and the perspective of end users. In this context, we observe a shift from well-curated and closed environments, e.g. museum exhibitions, guide tours, where users were exposed only to a pre-selected set of objects, carefully arranged in stories and narratives. Nowadays, the users have access to an endless pool of information, but the meaning of the individual objects is lost, because of often incomplete descriptions, lack of links between the objects and different collections.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Lora Aroyo is an IBM faculty award winner 2012 for Computer Science. She is an Associate professor for Web and Media at the VU Amsterdam and Scientific coordinator of EU Projects such as NoTube, for integration of Web and TV data with the help of semantics and CHIP, for Cultural Heritage Information Personalization. She is currently on at IBM Research. &lt;br /&gt;
Her research focuses on using semantic web technologies for modeling user interests and context and applying them in recommendation systems and personalized access to online cultural heritage collections, multimedia archives and interactive TV. She has coordinated the CHIP project on Cultural Heritage Information Personalization and the NoTube project on the integration of Web &amp;amp; TV with the help of semantics. She has co-organized numerous workshops on personalized access to cultural heritage, e-learning, interactive television, visual interfaces to the social and semantic web (e.g. PATCH, FutureTV, PersWeb, VISSW and DeRIVE). Lora is actively involved in the Semantic Web community, i.e. program co-chair for ESWC2009 and ISWC2011 and conference chair for ESWC2010. She is also actively involved in the Personalization and User modeling community as vice-president of UM Inc. and a member of the editorial board for the UMUAI journal.&lt;br /&gt;
&lt;br /&gt;
For more info: http://lora-aroyo.org&amp;lt;br&amp;gt;&lt;br /&gt;
For slides: http://www.slideshare.net/laroyo&amp;lt;br&amp;gt;&lt;br /&gt;
Follow me on twitter: @laroyo&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Tagasauris with Todd Carter==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Tagasauris the smart way to tag. Tagasauris is a start up based in New York City that develops innovative services and products .&lt;br /&gt;
&lt;br /&gt;
[[Todd Carter]] is the CEO and Co-founder of Tagasauris, a meta-data curation platform that makes your content smarter! We combine crowdsourcing, machine learning and semantic technologies that increases the lifetime value of digital media by making it more discoverable, connected and engaging. Todd is widely respected as a creative visionary and leader in the digital asset, photography and linked open data community, with over 20 years experience working with photo archives, libraries, museums and information technology systems. Tagasauris has been featured in The New York Times, Wired, Business Week, The Economist and others. The National Endowment for the Humanities awarded Tagasauris and The Museum of the City of New York a grant to annotate the museum&#039;s archive. Tagasauris is launching it&#039;s first photo tagging app that promises to be a game-changer in the media and entertainment industry. Tagasauris was founded in December 2010 with headquarters in New York City.&lt;br /&gt;
&lt;br /&gt;
Dealing with raw data at scale presents present a grand challenge. This is particularly true for visual media where the the large disparity between descriptions of multimedia content that can be computed automatically and the richness and subjectivity of semantics used to find, organize and share visual media present what is called the “semantic gap”.&lt;br /&gt;
One way to the address this gap would be to relabel raw data. Use experts. Describe the features of visual media with information about their content. But this is impossibly slow and prohibitively expensive! There has to be a better way.&lt;br /&gt;
Tagasauris has built a large-scale human computation engine that addresses the grand challenge posed by the semantic gap. The tagasauris process automates the discovery and generation of semantically linked metadata and enables us to quickly and cost effectively build a new, interlinked data layer of context for multimedia on the web that &amp;quot;fills in&amp;quot; the semantic gap.&lt;br /&gt;
The power of this layer will come from the quantity and quality of links between media objects on the visual web and their intersection other objects like events that can be detected and modeled from within the interest and social graphs. &lt;br /&gt;
We’re interested in discussion and exploration of uses cases that exploit the power of a personalized and semantically interlinked visual web.&lt;br /&gt;
For more information please visit www.tagasauris.com or email: info@tagasauris.com&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:Event]]&lt;/div&gt;</summary>
		<author><name>Marco</name></author>
	</entry>
</feed>