The Experts below are selected from a list of 321 Experts worldwide ranked by ideXlab platform

Henda Hajjami Ben Ghezala - One of the best experts on this subject based on the ideXlab platform.

  • Quality Based Data Integration for Enriching User Data Sources in Service Lakes
    2018 IEEE International Conference on Web Services (ICWS), 2018
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    Data lakes have recently emerged as an alternative solution to costly Traditional Data Warehouse solutions. To exploit Data lakes, however, there is a need for means that assist users in combining and integrating Data stored within a Data lake. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him/her. We show how Data lakes, or more specifically the service lakes, since we are focusing on Data providing services, can be leveraged to answer user queries, taking into account the quality of the services and respecting the (time and monetary) budget set by the user.

  • ICWS - Quality Based Data Integration for Enriching User Data Sources in Service Lakes
    2018 IEEE International Conference on Web Services (ICWS), 2018
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    Data lakes have recently emerged as an alternative solution to costly Traditional Data Warehouse solutions. To exploit Data lakes, however, there is a need for means that assist users in combining and integrating Data stored within a Data lake. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him/her. We show how Data lakes, or more specifically the service lakes, since we are focusing on Data providing services, can be leveraged to answer user queries, taking into account the quality of the services and respecting the (time and monetary) budget set by the user.

  • BIS - On Enriching User-Centered Data Integration Schemas in Service Lakes
    Business Information Systems, 2017
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    In the Big Data era, companies are moving away from Traditional Data-Warehouse solutions whereby expensive and time-consuming ETL (Extract-Transform-Load) processes are used, towards Data lakes, which can be viewed as storage repositories holding a vast amount of raw Data. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him. In this context, we show how the Data lake, or more specifically the service lake since we are focusing on Data providing services, can be leveraged to enrich the local Dataset with concepts that cater for the processing of user queries. Furthermore, we present the algorithms we have developed for this purpose and showcase the working of our solution using a study case.

  • On Enriching User-Centered Data Integration Schemas in Service Lakes
    Business Information Systems, 2017
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    In the Big Data era, companies are moving away from Traditional Data-Warehouse solutions whereby expensive and time-consuming ETL (Extract-Transform-Load) processes are used, towards Data lakes, which can be viewed as storage repositories holding a vast amount of raw Data. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him. In this context, we show how the Data lake, or more specifically the service lake since we are focusing on Data providing services, can be leveraged to enrich the local Dataset with concepts that cater for the processing of user queries. Furthermore, we present the algorithms we have developed for this purpose and showcase the working of our solution using a study case.

Pedro Furtado - One of the best experts on this subject based on the ideXlab platform.

  • near real time with Traditional Data Warehouse architectures factors and how to
    International Database Engineering and Applications Symposium, 2013
    Co-Authors: Nickerson Ferreira, Pedro Martins, Pedro Furtado
    Abstract:

    Traditional Data Warehouses integrate new Data during lengthy offline periods, with indexes being dropped and rebuilt for efficiency reasons. There is the idea that these and other factors make them unfit for realtime warehousing. We analyze how a set of factors influence near-realtime and frequent loading capabilities, and what can be done to improve near-realtime capacity using a Traditional architecture. We analyze how the query workload affects and is affected by the ETL process and the influence of factors such as the type of load strategy, the size of the load Data, indexing, integrity constraints, refresh activity over summary Data, and fact table partitioning. We evaluate the factors experimentally and show that partitioning is an important factor to deliver near-realtime capacity.

  • IDEAS - Near real-time with Traditional Data Warehouse architectures: factors and how-to
    Proceedings of the 17th International Database Engineering & Applications Symposium on - IDEAS '13, 2013
    Co-Authors: Nickerson Ferreira, Pedro Martins, Pedro Furtado
    Abstract:

    Traditional Data Warehouses integrate new Data during lengthy offline periods, with indexes being dropped and rebuilt for efficiency reasons. There is the idea that these and other factors make them unfit for realtime warehousing. We analyze how a set of factors influence near-realtime and frequent loading capabilities, and what can be done to improve near-realtime capacity using a Traditional architecture. We analyze how the query workload affects and is affected by the ETL process and the influence of factors such as the type of load strategy, the size of the load Data, indexing, integrity constraints, refresh activity over summary Data, and fact table partitioning. We evaluate the factors experimentally and show that partitioning is an important factor to deliver near-realtime capacity.

Hiba Alili - One of the best experts on this subject based on the ideXlab platform.

  • Quality Based Data Integration for Enriching User Data Sources in Service Lakes
    2018 IEEE International Conference on Web Services (ICWS), 2018
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    Data lakes have recently emerged as an alternative solution to costly Traditional Data Warehouse solutions. To exploit Data lakes, however, there is a need for means that assist users in combining and integrating Data stored within a Data lake. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him/her. We show how Data lakes, or more specifically the service lakes, since we are focusing on Data providing services, can be leveraged to answer user queries, taking into account the quality of the services and respecting the (time and monetary) budget set by the user.

  • ICWS - Quality Based Data Integration for Enriching User Data Sources in Service Lakes
    2018 IEEE International Conference on Web Services (ICWS), 2018
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    Data lakes have recently emerged as an alternative solution to costly Traditional Data Warehouse solutions. To exploit Data lakes, however, there is a need for means that assist users in combining and integrating Data stored within a Data lake. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him/her. We show how Data lakes, or more specifically the service lakes, since we are focusing on Data providing services, can be leveraged to answer user queries, taking into account the quality of the services and respecting the (time and monetary) budget set by the user.

  • BIS - On Enriching User-Centered Data Integration Schemas in Service Lakes
    Business Information Systems, 2017
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    In the Big Data era, companies are moving away from Traditional Data-Warehouse solutions whereby expensive and time-consuming ETL (Extract-Transform-Load) processes are used, towards Data lakes, which can be viewed as storage repositories holding a vast amount of raw Data. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him. In this context, we show how the Data lake, or more specifically the service lake since we are focusing on Data providing services, can be leveraged to enrich the local Dataset with concepts that cater for the processing of user queries. Furthermore, we present the algorithms we have developed for this purpose and showcase the working of our solution using a study case.

  • On Enriching User-Centered Data Integration Schemas in Service Lakes
    Business Information Systems, 2017
    Co-Authors: Hiba Alili, Rim Drira, Daniela Grigori, Khalid Belhajjame, Henda Hajjami Ben Ghezala
    Abstract:

    In the Big Data era, companies are moving away from Traditional Data-Warehouse solutions whereby expensive and time-consuming ETL (Extract-Transform-Load) processes are used, towards Data lakes, which can be viewed as storage repositories holding a vast amount of raw Data. In this paper, we position ourselves in the recurrent context where a user has a local Dataset that is not sufficient for processing the queries that are of interest to him. In this context, we show how the Data lake, or more specifically the service lake since we are focusing on Data providing services, can be leveraged to enrich the local Dataset with concepts that cater for the processing of user queries. Furthermore, we present the algorithms we have developed for this purpose and showcase the working of our solution using a study case.

Xiaofang Li - One of the best experts on this subject based on the ideXlab platform.

  • ICIA - Real-Time Data ETL framework for big real-time Data analysis
    2015 IEEE International Conference on Information and Automation, 2015
    Co-Authors: Xiaofang Li
    Abstract:

    In the big Data era, Data become more important for BI and SCADA system operation. The load cycle of Traditional Data Warehouse is fix and longer, which cannot timely response the rapid Data change. Real-time Data Warehouse technology, as an extension of Traditional Data Warehouse, can capture the rapid Data change and process the real-time Data analysis to meet the requirements of SCADA system. The real-time Data access without the processing delay is a challenging task to the real-time Data Warehouse. In this paper, the real-time Data ETL framework is presented to separately process the historical Data and real-time Data. Then, combining an external dynamic storage area, a dynamic mirror replication technology was proposed to avoid the contention between OLAP queries and OLTP updates. Finally, the experiments is set up based on the TPC-H benchmark to evaluate the performance of the proposed real-time Data ETL framework. The experimental results demonstrates the proposed solution to real-time Data ETL can effectively mitigate the query contention and Data skew.

  • Real-Time Data ETL framework for big real-time Data analysis
    Information and Automation, 2015 IEEE International Conference on, 2015
    Co-Authors: Xiaofang Li, Yingchi Mao
    Abstract:

    In the big Data era, Data become more important for BI and SCADA system operation. The load cycle of Traditional Data Warehouse is fix and longer, which cannot timely response the rapid Data change. Real-time Data Warehouse technology, as an extension of Traditional Data Warehouse, can capture the rapid Data change and process the real-time Data analysis to meet the requirements of SCADA system. The real-time Data access without the processing delay is a challenging task to the real-time Data Warehouse. In this paper, the real-time Data ETL framework is presented to separately process the historical Data and real-time Data. Then, combining an external dynamic storage area, a dynamic mirror replication technology was proposed to avoid the contention between OLAP queries and OLTP updates. Finally, the experiments is set up based on the TPC-H benchmark to evaluate the performance of the proposed real-time Data ETL framework. The experimental results demonstrates the proposed solution to real-time Data ETL can effectively mitigate the query contention and Data skew.

Nickerson Ferreira - One of the best experts on this subject based on the ideXlab platform.

  • near real time with Traditional Data Warehouse architectures factors and how to
    International Database Engineering and Applications Symposium, 2013
    Co-Authors: Nickerson Ferreira, Pedro Martins, Pedro Furtado
    Abstract:

    Traditional Data Warehouses integrate new Data during lengthy offline periods, with indexes being dropped and rebuilt for efficiency reasons. There is the idea that these and other factors make them unfit for realtime warehousing. We analyze how a set of factors influence near-realtime and frequent loading capabilities, and what can be done to improve near-realtime capacity using a Traditional architecture. We analyze how the query workload affects and is affected by the ETL process and the influence of factors such as the type of load strategy, the size of the load Data, indexing, integrity constraints, refresh activity over summary Data, and fact table partitioning. We evaluate the factors experimentally and show that partitioning is an important factor to deliver near-realtime capacity.

  • IDEAS - Near real-time with Traditional Data Warehouse architectures: factors and how-to
    Proceedings of the 17th International Database Engineering & Applications Symposium on - IDEAS '13, 2013
    Co-Authors: Nickerson Ferreira, Pedro Martins, Pedro Furtado
    Abstract:

    Traditional Data Warehouses integrate new Data during lengthy offline periods, with indexes being dropped and rebuilt for efficiency reasons. There is the idea that these and other factors make them unfit for realtime warehousing. We analyze how a set of factors influence near-realtime and frequent loading capabilities, and what can be done to improve near-realtime capacity using a Traditional architecture. We analyze how the query workload affects and is affected by the ETL process and the influence of factors such as the type of load strategy, the size of the load Data, indexing, integrity constraints, refresh activity over summary Data, and fact table partitioning. We evaluate the factors experimentally and show that partitioning is an important factor to deliver near-realtime capacity.