quinta-feira, 16 de maio de 2024

How Data Analytics and Data Science Fit: A Join Research Methodological Perspective - By Faiza Bukhsh and Maya Daneva

Definition: Design Science is the design and inestigation of artifacts in context.

Design - design/investigate scinticif methods processes, algorithms.
Investigation - investigat from noisy, structure and unstructured data
Artifacts - extract or extrapolate knowledge and insights
Context - apply in context the knowlede grom data across a broad range of application domains.

There are different perspectives to define/use Data Science, depending on your field. See figure below:
The basic phases of Data Analytics can also be seen as the basic phases of Design Science (it really depends on the project how to frame each phase. For instance, for a software, the phases have to do with software development)
CRISP-DM: a method of Data Mining, which proposes a life cycle which is similar to any design project. It comprehends Business Understandingg, Dta Understanding, Data Preparation, Data Modeling, Evaluation and Deployment. It is an iterative and cyclic process. See figure below:
Different roles are responsible for each activity but they should also collaborate and participate in each other activities. Example:
Many times, the problem happens because the data is not well-prepared. We get excited to run our model and train it, but if the data is noisy, not well-prepared, you will never achieve the accuracy that you are looking for. So you should stop doing mindless effort and go back to data preparation. For acquiring data, there are a few possibilities. One solution she made with a hospital is having the data aways on their site on a particular server and then she can access it through a VPN. But for privacy issues, she cannot access the data itself, only the results of the application of algorithms. *This seems like a promising solution There is also another interesting methodology which has the same phases, but what is special about it is that it always loops back to a previous activity. See figure:
There is a very interesting method called SEMMA. Sample, Explore, Modify, Model and Assess. They have a setp-by-step guidance of how to tackle each phase:
She mentions four methodologies for Design Science. Among them, Roel Wieringa's, Paul Johanssen's and Peffer's. let's start with Paul Johanssen's method
Paul's model has a lot of cycles on it. The Data Science methods also have cylces. So we can start analyzing in which parts of one model we can insert the other. Peffers Design Cycle
Wieringa's method:
In the work of Wieringa, data analysis will be in the treatment design, since this is where you do the modeling, the training and the tunning. But if you are going for a knowledge problem, then it means that you are trying to extract knowledge from knowledge (so something like LLM). Then, data analysis is in the setup phase (for LLM, it will be prompt engineering). How can I know what my artifact is? (the artifact that should be designed)
It all depends on the objective. You have to ask yourself what is the goal of that design project.

What is my artifact?
The artifact in a Data Analytics project can be the Data preparation process itself, it can be the model or the model result, it can be the evaluation process or evaluation criteria. So we have to ask ourselves again what the objective is. And then you will know what is your artifact.

Keynote by Carlos Ribas (Bosch) - The power of Information Systems shaping the future of the Automative Industry

He presented dan interesting tool called Bmlp associated with an operating system named TOM to automate smart factories. Read about it here: https://www.iotm2mcouncil.org/iot-library/news/connected-industries-news/bosch-commits-to-global-industrial-aiot/

He also discussed how Bosch inveted in Digital Twins to help having more prompt predictions of failures in their factories. Read about it here: https://www.bosch-connected-industry.com/de/en/iiot-insights/digital-twins

He talked about the role of AI in Manufacturing
Examples of use:
He also mentioned that none of this is important if tecnhology does not improve the life or work of people who work in the factories.
Digital twins can help simulate in the lab before the equipment is put in the plant. The other use is in real-time data acquisition. For example, when the equipment is put to test, at the same time, they can inspect data coming from the test and discover on the flight. And they will know that in a particular component, under specific conditions, they have errors. This is really helpful for them.

They treat data integration in these terms: from each sensor, data is sent to data repositories in specific format and also adding labels. This facilitates recover data from different apps, different systems. *It seems to me they treat this in the syntactic level.

I also found a link to an interesting data platform: https://www.bosch-connected-industry.com/de/en/portfolio/bosch-semantic-stack
I wonder if there are some more sophisticated semantic technologies in place, of which perhaps Carlos is not aware.

They can detect a problem in the process, not after the process is finished. That is why they feel so much in control. The faulty components are immediately rejected, removed from the process, suffer maintenance, and then go back to the process.
They are not currently investing in LLM because they do not feel the need. Sometimes the volume of data being too high, it does not help.

The people who used to work in the plant doing mechanical work are still there, and they are trained and "re-skilled". In the last years, the process of training has been very intensive. Sometimes, they are not learning new things very easily, but Bosch sees this is a mission. If they do not

People need to develop different competencies. In the future, in the recruitment process, new people need to come with a degree. Currently, many of them are low level engineers (now the work force is 40% of people have at least a degree). They need to gain knowledge about the new technologies. This is a must!

Alessandro Oltramani, an expert in logic-symbolic reasoning is the new leader of the Carnegie Bosch Institute: https://carnegiebosch.cmu.edu/

Giancarlo asked if this shows that such kind of approach is a current bet of Bosch. Carlos responded that is for sure.

quarta-feira, 3 de abril de 2024

What are the Ontological Foundations of Simulation Modeling - Gerd Wagner - SCS Weekly Meeting

What are the Ontological Foundations of Simulation Modeling In simulation modeling, you care about modeling objects and events, since we want to simulate the real world, and these are the most important types of ontological categories in the world. Modeling and simulation (M$S) is concerned with modeling dynamical systems which considt of ojbect that are subject to state changes over tim. This happens when one or more of its attribute values are changed. These attributes that change are called state variables. Attribute values may be continuous (smooth) or discrete (in jumps), leading to continuous or discrete processes. In ISs we are typically more concerned with discrete processes. Sometimes, a mix of them: discrete events (e.g. a car bumping into another in traffic) but also continuous events (movement of different objects in traffic). Discrete Systems: - Example: predator-prey ecosystem such as an area populated by wolves and sheep, where births, death and predator-prey encounters are events. - A discrete dynamical system (such as the one in the example above) can be captured either more abstractly with the help of a continous simulation as in System Dynamics, or iwth the help of a Discrete Event Simulation model. Discrete Evnet Simulation (DES) Paradigms: - Event-based simulation with SIMSCRIPT (1962), Event Graphs (1983) - Process Network simulation with GPSS (1961), Arena (1992), AnyLogic etc. It is based on more high-level concepts w.r.t events, which help you to capture concepts of different domains (e.g. manufacturing, traffic etc.) - Coroutine-based Process Interaction simulation with Simla (1967), SimPy, etc. Coroutines are asynchronous programming process stations, which may start, be interrupted and then reestablish processing. - Simulation based on Petri Nets (from the 60s) Object Event M&S Based on the ontological principles: - objects participate in events - events cause state changes of participating objects and follow-up events according to causal regularities. The sturcture of objects and events is described in the form of a UMLclass mode defining object types and event types. The system's dynamics is described in the forms of DPMN (similar to BPMN) process model defining a set of rules. - which caputre causal regularities (as event rules) - and correspond to transition functions of an Abstract State Machine. Causal Regularity Simple Model:
Example of Object Event Model about Phishing We may see an OE Class Model to model the information and a BPMN/DPMN model to model the events
Agent-based M&S - ontologically speaking, agents are special objets that interact with each other and with their environment. - agents interact with their envionment via a perception-action cycel that is modeled in OEM&S in the form of perception events and action events. - Agents interact with each other by sending and receiving messages. In OEM%S, sending a message is an out-message action event and reeiving a message is an in-message event. Example of a basic BPMN Model of Phishing
Example of a Conceptual Information Model about Phishing
Besides the regular relationships (composition, specialization, and general associations), in these kind of Information Models, there are special kinds of associations and multiplicity restrictions: - the association between the entities mean the participation of agents/objects in events. - the multiplicity can indicate snapshot or historical multiplicity restrictions (you may need both kinds of multiplicity in one model).

sexta-feira, 8 de março de 2024

Crafting Future Scenarios with the Help of AI - Roland M. Mueller, Katja Thoring at al.

Developing future research poses some problems, including the fact that you don't have the users for manymuch the technology you want to produce. Research questions: Goals: provide AI assistance to future scenario development; AI assistance with scenario rating; AI assisatance with Qualitative Feedback, AI assistance with scenario iterations etc. - Can we democratize access to collective expert knoweldge through Generative AI? - Can we expand the established - Can we build human twins to Project called: Delphi Study Experiment 1: They developed 23 future scenarios using a panel of experts: people from different non-AI fields, such as science fiction authors, business people. And they conmpared that with the ideas of the people in the AI research field. E.g. of solution of the painel of experts: Digital Detox Zone (a place in the office which is not digitally supported) Experiment 2: compare the Human expers and AI experts with a Digital Twin. In short, it does not work yet. Paper to read: Designing the Future With the “Delphi Design Sprint”: Introducing a Novel Method for Design Science Research - https://www.researchgate.net/publication/357746370_Designing_the_Future_With_the_Delphi_Design_Sprint_Introducing_a_Novel_Method_for_Design_Science_Research Discussion about the use of Digital Twins in these scenarios: - Good potential for triangulation with field experts and AI people. - Good inspiration for future works in this area Ethical concerns: - GenAI hallucinations are not asuch aproblem for scenairo development compared to factual quesitons - Tranparency of AI involvement - Specific requirement and charactiristic of AI scneaqrio Crafting Future Scnarios with the Help of AI: Potentials of a Hybrid Delphi Expert Panel. HICSS Mind th eFuturee Gap: Introducting the FOD Framework for Future Oriented Design. HICSS

Digital Everything: From Twins to Circular Economy - Barbara Dinter

Digital Everything: From Twins to Circular Economy Barbara Dinter Barbara is one of the IS chairs, focusing on Business Intelligence in TU Chemnitz This presentation is about some german-funded projects. Project 1 - Co-Twin - Vision of a collaboration digital tiwn (DT) in value chain networks. - Whole life cycle - she applies Business Models They transfered the ARIS idea of views (BP view, Data view etc.) to Digital Twins. They have: component view, data view, visualization view, network view... and others. - For all stakeholders in a value chain. The DT is used on the planning phase Results: demonstrator prottoype, 3 use cases, taxanmy, reference architecture, design guidelines and conceptualizations. Project 2 - The circular economy - Part 1 - integration with Co-Twin project Goals: sustainability, enrionmental protection and increased efficiency. Key aspects: Reduce, reuse, repair and recycle; sustainable business models, systemic approach, design for longevity and integration of digital technologies. - Part 2 - The circular economy Digital ecosytem for circular economy in the automotive industry (DIONA) collaboration with other academic partners: TU Dortmun and Fraunshofer ISST. She also mentioned 12 projects with industry. DIONA Focus areas - Transfer and networking: coordination of 12 MobilKreis projects, oraganization of physical and digital meetings, knwoeldge tranfer research activities. - Cyberphysical Lab for SMEs to open experimental space for simulations and test and vailidate scnearios without disrupting live processes. - Digital Hub Research topics: 1) conceptualization of use caes in Circular economy 2) BPM in Circular Economy (adaptation of capabilities, models and technologies for that)

segunda-feira, 3 de julho de 2023

Fundamentals of Disaster Management Sytems: A Computer Science Perspective - Mehmet Aksit - University of Twente - 3-7-2023

Fundamentals of Disaster Management Sytems: A Computer Science Perspective Mehmet Aksit - University of Twente - 3-7-2023 Outline: Disaster Management as it is today Problem Statement Process Automation for Disaster Management Today: 1) Disaster Prediction
2) Post-Disaster:
- Identification
- Assess
- Understand
- Cope
- Strategy
- Recovery Procedures

Problem statement He worked with the auhtorities and developed requirements for dealing with disaster and without looking at the trends, really focused on talking to authoraties' stakeholders. They used different viewpoints to elicit requirements
They used quesionnaires with wanted-not wanted 5 likert scale in a large organization in Turkey. This study identified 65 requirements that are now prioritized. For prioritization, they identified which ones could be deffered to a later time without causing technical problems during disaster management (priorization in 3 groups or requirements). After that, they synthesized to identify technologies and skills to fulfill the requirements.
Obsserved problems: they analyzed the problems and saw that it would be difficulty to deliver Process Automation in the way they intended.
They generalized the problem of process automation as a problem of Demand and Supplies. There are several related works in aid optimization for disasters. But they have some limitations: generally offline optimizations, problem-specific (case by case), product specific (for specific systems), weak or absence of automated process support, accordingy, lcak of an online control system platform to manage disasters effectively and efficiently.

Solutions for the demand and supply problem
Example: They have about 100 tables like the one in the picture below, with established rules. Then the cocneptual model on the right hand corner guides the automation of a process to deal with a particular disaster. They create tasks and group tasks that compose jobs, and this is assigned to different people.
If there isn't enough resource, then some strategies are developed (trade-off analysis, prioritization of jobs, etc.)
To sum up, for them, a disaster problem is a resource allocation problem.
Upcoming publication:



Disaster Prediction Another direction of work they do is to predict disasters. They do that based on the concept of 'Event'.
I asked them if they do that by analyzing social media, but he said 'no'. They use other data sources.
Giancarlo has asked an interesting question about causality, in other words, how to identify that an event will happen because it is caused by a previous event occurence that has been already identified. Mehmet said that they treat this in their approach.
Event specification:
Upcoming publication:
Using Digital-Twins for disaster prediction and handling:
Conclusions
From a ML point of view, there is a lot to be developed such as:
Notes on the discussion subsequent to the talk:

There are many advanced simulation systems out there, and instead of producing new ones, we should use the existing systems that are already very sophisticated.

They keep collaborations with the Disaster Center at the Univ. of Sao Carlos, in Brazil, and they hope to talk to the Federal Disaster Management team as well. The main problem that he sees is that Brazil, like in Turkey and many other countries, there is a lot of data from satelites and other state of the art sensors, but there is no data on disaster management process. And the main problems for him are process problems. That is why they are working on process automation, and these partners have been very excited about this kind of work.

It is even relatively easy to find data sets on disasters, but not on disaster management processes. Therefore, it is very important to work on producing this data, as well as integrating such data from different data sources. Also during the pandemic he noticed that most of the work was on putting data together from different systems. That is why we need International Alliances and a lot of work on data integration.

Their system is responsible to make decisions to allocate tasks to teams in disaster management. The decisions are based on rules and ML algorithms that adapt the rules for making the predictions and recommendations. How to cope with people's resistance in using such systems? Jan says that perhaps the operator should give the last word and not the system. But actually in many areas, we see systems that already take the decision on our behalf and they can do it better than us (e.g., piloting planes, flood gates in Rotterdam). Mehmet says that something important is to categorize the type of disaster - meaningful categories! (see slide on ML opportunities above). Depending on the type of disaster, there can be more or less human intervention. Giancarlo suggests that work should be done in: intentional (goal) modeling to understand what people's intention is in the middle of a disaster (e.g. Paris riots). Mehmet says: another important issue is to deal with ethics and privacy in this domain.

segunda-feira, 16 de dezembro de 2019

Unitn Class - 16-12-2019 - Seminar on AlpineBits - By Claudenir Morais Fonseca

AlpineBits
By Claudenir

Focus: create a standard for touristic events of the Alto Adige region.

The standard is based on formats such as XML/JSON

The standard is free and may be downloaded from the AlpineBits website.

It is not a service... it is a contract to use a particular format on data exchange.

Main motivation: to exchange knowledge in a common format,, so that all businesses and the government can consume this data. In the end, the businesses are competitors, but they collaborate on developing the standard, because it is advantageous for all of them.

------------------------------

Scope

Touristic events: events, event series, venues, organizers etc.
Touristic areas: lifts, trails,, slopes, mountain areas etc.
Additional information: agents, multimedia, metadata etc.

-------------------------------

Eventdata reference model

*He showed several models, some were class models, other instance models. This was very interesting to illustrate the usefulness of the approach.

The more complete model contains the following concepts: event, event series, venue, venue allocation, the roles of the people and organizations connected to the event, and target audience.

--------------------------------

They reuse as much as possible. For example, they use schema.org.
Reusing is important to guaranteeing acceptance by partners (people want to use popular solutions, which they are already familiar with).
However using the things they reuse alone is not enough, because this would keep the information too vague.
So their solution builds on top of the existing artefacts/tools

-------------------------------

They applied GitLab and created issues for everything they had to do for the project.
And they followed the good practice of having a structured development, using features path, development path and master path, leaving the master path as clean as possible.

The repository is public and thus available for the students to take a look

----------------------------------

One of the most important things in Claudenir's view is testing. You must have a way to guarantee the quality of the code, especially if you expect someone else to revise your code.

Here is the testing tools Alpine Bits make available.

--------------------------------------

Some requirements
minimize message size
one way to respresent information
(...)

-----------------------------------

Main use cases

  • Request-Response
  • Publisher-Subscriber
  • Batch exchanges
-----------------------------------

Interaction Types

  • Request-Response
  • Publisher-Subscriber
  • Batch
  • Streaming
------------------------------------

API Styles

  • Tunnel Style - having one endpoint and everything else goes on the message
  • URI Style (Partial REST)
  • Hypermedia (Full REST)
  • Query Style
  • Event-driven Style
--------------------------------------

In general, the resources for AlpineBits developers are found here.