Saturday, October 10, 2015

Conflation in GIS

Conflation is the action of unifying two distinct datasets into a new dataset. This may be relatively easy to extremely difficult depending upon the complexity of representation and the size and quality of datasets involved.

Conflation terminology:

'Matching' is the activity of identifying features or data elements that represent the same real-world entity.

'Alignment' describes the degree to which the two features or data elements have coincident geometry

'Adjustment' is the alteration of geometry or attributes of matched features to align them.

The 'reference dataset' is the one which is to be conflated to. It is of greater spatial accuracy than the 'subject dataset'

The 'subject dataset' is the dataset to be matched or adjusted.

Conflation problems can be classified into:
-Horizontal
-Vertical and
-Internal

Horizontal conflation is the process of eliminating discrepancies along the common boundary of datasets that are adjacent to one another. These include datasets containing data from same feature classes. For example, aligning boundaries of adjacent coverages or edge-matching neighbouring networks.

Vertical conflation involves matching or eliminating discrepancies between datasets that occupy the same area in space. For example, road network matching between two representations of roads in the same region.
Two important types of vertical conflation are:
-Version matching and
-Feature Alignment

In version matching, the input datasets consist of different versions of the same features. The conflation process helps identify matching features. Attributes are transferred between matched features and unmatched features are transferred completely. For example, matching different versions of road networks for the same geographical area.

In feature alignment, the input data consists of features from two or more different feature classes that have some defined relationship to each other. The conflation process is aimed at removing discrepancies between datasets that falsify this relationship. For example, geometric alignment. A specific example in this context is that of aligning boundaries of different kinds of feature classes such as municipal districts

Internal conflation involves resolving features or element within a single dataset. For example, coverage cleaning may require removal of overlaps in polygons within a coverage.

CONFLATION WORKFLOW:
The process of conflation can be broken down into the following sub tasks:
i. Data pre-processing: This step normalizes the datasets and ensures that they are compatible. This may involve format translation and other basic preparation of the datasets. An example of data pre-processing is to ensure that the datasets must have the same coordinate system.
ii. Data quality assurance: In this step, the internal consistency of the datasets is verified and improved if necessary. Sometimes, conflation tasks require that datasets have an internal level of consistency. For example, coverage alignment algorithms require that the input datasets are a clean coverage.
iii. Dataset alignment: In case the datasets are mis-aligned, an initial alignment process is required to carry-out precise conflation. This alignment is coarse-grained in nature and does not align individual features.
iv. Feature matching: This step involves matching of common features between datasets. After this phase is performed, the discrepancies between datasets would have been identified and can be visualised. It is used to provide statistical summaries of data quality.
v. Geometry alignment removes discrepancies between geometries
vi. Information transfer involves updating one dataset with information from the other. This information can be either attributes or geometry to be added to an existing feature or entire features to be added to the dataset.

SCHEMATIC REPRESENTATION OF CONFLATION WORKFLOW


Monday, October 5, 2015

Object Structural Model in GIS

DBMS (Data Base Management System) play an important role in GIS by forming a link between spatial and aspatial data. DBMS provide access to data to multiple users at the same time, prevent loss of data and provide security access. Currently available DBMS are generally based on the hierarchical model,  network model, relational model or their derivatives. Columns of the table are called attributes and all the values of an attribute describe the set of all possible values. Rows are also called records, tuples or relation elements.
However, the data model on which most relational DBMS are based do not meet the requirements of modern GIS. GIS integrate data from a variety of sources into a single homogeneous system and hence need powerful and flexible data models to meet the requirements of multiple tasks. For example:

  1. Sophisticated treatment of real world geometry
  2. Representation of data at different conceptual levels of resolution and detail
  3. management of history and versions of objects
  4. Combinations of measurements of different resolutions and accuracy

Remotely sensed data

GIS and Remote Sensing are linked by remotely sensed data, mostly in the form of aerial or satellite images of a specified piece of land are used as input into GIS for manipulation and analysis. This in turn is presented to policy makers for the formulation of policies in the context of development or conservation of resources.

For many applications, remote sensing can be used effectively and efficiently to update GIS data layers. These updated layers in GIS can be used to improve the interpretability and information extraction potential of remotely sensed data.

GIS users require timely input data to optimise their systems for analysis and decision making. Usually, raster data is preferred. GIS users require multiple multiple images from different regions of the electromagnetic spectrum and/or different dates at scales ranging from local to global. Moreover, the data might have a variety of spatial, spectral and temporal resolutions.

Maps are compiled using photogrammetric techniques to process remotely-sensed data and these maps form the base upon which GIS applications are achieved. Remotely sensed data are also used to measure several environmental parameters like: surface and cloud top reflectances, albedo, soil and snow water content, fraction of photosynthetically active radiation, areas and potential yield of crop types, height and density of forest stands, etc. Such data mapped and/or monitored over time form the basis for monitoring.

Environmental planners, resource managers and public-policy decision makers are employing remotely-sensed data within the context of GIS to improve the management of resources.

Remote sensing data are acquired by aerial camera systems and a variety of active and passive remote  sensor systems operating at wavelengths throughout the electromagnetic spectrum. Data acquired by aerial camera systems can be scanned, converted into digital format and input into GIS.

Most satellite scanners are typically electro-mechanical scanners, linear devices or imaging spectrometers that operate in either a 'sweep' (LANDSAT) or 'pushbroom' (SPOT) mode. These are passive systems that record solar radiation reflected from Earth's surface.

Data derived from multi-spectral scanners can provide information on, vegetation types, its distribution and condition, geomorphology, soils, surface waters and river networks.

Short Wave Infra-Red (SWIR) sensors record emitted energy from surfaces and have been particularly useful for monitoring fires and studying areas of geothermal and volcanic activity. Thermal or Long Wave Infra Red (LWIR) sensors are used for mapping ocean temperatures and study of the dynamics of ocean waters and currents. Thermal maps are used to monitor urban areas, industrial sites, manufacturing centers and agricultural fields.

Active systems operating in the visible spectrum use laser technologies (Ex: LIght Detection And Ranging Systems (LIDARS)) mainly for oceanographic and forestry applications. Regardless of the wavelengths they use, active systems DO NOT depend on the sun for image capture.

The following is a list of satellite based sensors currently providing operational raster remotely sensed data for GIS developers and users:

  1. LANDSAT programme: -
      1. It has provided coverage of Earth for almost 25 years
        1. It is the result of NASA Earth Resources Survey Program and several other U.S. government agencies
        2. It was originally known as Earth Resources Technology Satellite in 1972.
        3. Four additional satellites have been placed in orbit since 1972 for providing continuous data for use in a wide range of environmental applications
        4. The first three LANDSATs had a Multi Spectral Scanner (MSS) as the primary sensor while the next two had a high resolution scanner called the Thematic Mapper (TM)
        5. MSS had a spatial resolution of 80m and images in the visible and near-infra Red region while the TM had a resolution of 30m and images in the visible and thermal infra-Red band.
        6. The LANDSAT programme established the operational viability of space-based remotely sensed data
    1. Satellite Pour I'Observation de la Terre:-
        1. SPOT is an operational, commercial remote sensing programme that operates on an international scale. 
        2. SPOT satellites are owned and operated by a french space agency, Centre National d'Etudes Spatiales (CNES). 
        3. Three SPOT satellites have been placed into orbit since 1986.
        4. The mission objectives for SPOT are:
          1. providing remotely-sensed data suited for land cover, agriculture, forestry, geology, regional planning and cartography applications.
        5. Data from the High Resolution Visible sensor (HRV) provide both multispectral coverage with 20m spatial resolution and panchromatic imagery with 10m resolution. This data is particularly well suited for urban and cartographic applications.
    2. Advanced Very High Resolution Radiometer:-
        1. The AVHRR sensor is carried aboard the USNOAA's (United States National Oceanic and Atmospheric Administration's) Polar Orbiting Environmental Satellites (POES)
        2. This program was established to provide data for use in meteorological applications.
        3. The daily coverage provided by AVHRR has resulted in the data being used for many operational land mapping and monitoring programs
        4. AVHRR data are multispectral and the data have a resolution of 1.1km at nadir and an orbital swath of 2600 km.
    3. Marine Observation Satellite
    4. Japanese Earth Resources Satellite
    5. India Remote Sensing Satellite
    6. European Resource Satellite
    7. RADARSAT
    8. High Spatial Resolution Satellites

    Conversion of existing digital data

    An important technique that is becoming increasingly popular for data input is the conversion of existing digital data. A variety of spatial data, including digital maps, are openly available from a wide range of government and private sources. The most common digital data to be used in a GIS is data from CAD systems. A number of data conversion programs exist, mostly from GIS software vendors, to transform data from CAD formats to a raster or topological GIS data format. Several adhoc standards for data exchange have been established in the market place. These are supplemented by a number of government distribution formats that have been developed. Due to the wide variety of data formats, most GIS vendors have developed and provide data exchange/conversion software to go from their format to those considered common in the market place.

    Most GIS software vendors also provide an ASCII (American Standard Code for Information Interchange) data exchange format specific to their product, and a programming subroutine library that allows users to write their own data conversion routines to fulfil their own specific needs. As digital data becomes more readily available this capability becomes a necessity for any GIS. Data conversion from existing digital data is not a problem for most technical persons in the GIS field. However, for smaller GIS installations who have limited access to a GIS analyst this can be a major problem in getting a GIS operational. Government agencies are usually a good source for technical information on data conversion requirements.

    Some of the data formats common to the GIS marketplace are listed below.

    IGDS - Interactive Graphics Design Software (Intergraph / Microstation)

    This binary format is a standard in the turnkey CAD market and has become a de facto standard in the mapping industry. It is a proprietary format, however most GIS software vendors provide DGN translators.

    DLG - Digital Line Graph (US Geological Survey)

    This ASCII format is used by the USGS as a distribution standard and consequently is well utilized in the United States. It is not extensively used even though most software vendors provide two way conversion to DLG.

    DXF - Drawing Exchange Format (Autocad)

    This ASCII format is used primarily to convert to/from the Autocad drawing format and is a standard in the engineering discipline. Most GIS software vendors provide a DXF translator.

    GENERATE - ARC/INFO Graphic Exchange Format

    A generic ASCII format for spatial data used by the ARC/INFO software to accommodate generic spatial data.

    EXPORT - ARC/INFO Export Format .

    An exchange format that includes both graphic and attribute data. This format is intended for transferring ARC/INFO data from one hardware platform, or site, to another. It is also often used for archiving.

    ARC/INFO data. This is not a published data format, however some GIS and desktop mapping vendors provide translators. EXPORT format can come in either uncompressed, partially compressed, or fully compressed format

    A wide variety of other vendor specific data formats exist within the mapping and GIS industry. In particular, most GIS software vendors have their own proprietary formats. However, almost all provide data conversion to/from the above formats. As well, most GIS software vendors will develop data conversion programs dependant on specific requests by customers. Potential purchasers of commercial GIS packages should determine and clearly identify their data conversion needs, prior to purchase, to the software vendor.

    Cartographic database

    Cartographic database:
    A database containing x-y coordinates defining a geographical area. When combined with other data (such as any of a wide variety of variables, such as income distribution, age, etc.), a cartographic database can be used to map the distribution of that variable within a geographical region.


    Digital Elevation Data

    Digital Elevation Data (DED) consists of an ordered array of ground elevations at regularly spaced intervals. The digital data for DED is extracted from the hypsographic and hydrographic elements of the various scaled positioned data acquired from the region in question.

    Data compression

    Data compression:
    Data compression in GIS refers to the compression of geospatial data so that the volume of data transmitted across networks can be reduced. A properly choosen compression algorithm can reduce data size upto 5 - 10% of the original image and 10 - 20% for vector and text data. Such compression ratios could result in significant performance improvement.
    Data compression algorithms can be classified into lossless and lossy . Lossless compression algorithms are those where the bit streams can be recovered to original data. These algorithms should be used if loss of even a single bit may cause serious and unpredictable consequences.
    Lossy compression algorithms should be used where a certain level of distortion can be tolerated. They are used to achieve a higher level of compression.
    Examples of lossless compression algorithms are:
    -Huffman coding
    -Arithmetic coding
    -Lempel-Ziw Coding (LZC) and
    -Burrows-Wheeler Transform (BWT)

    Examples of lossy compression algorithms are:
    -Differential Pulse Coded Modulation (DPCM)
    -Transform Coding
    -Subband Coding and
    -Vector Quantization

    Data compression refers to the process of reducing the size of a file or database. Compression improves data handling, storage, and database performance.
    Examples of compression methods include quadtrees, run-length encoding, and wavelets.

    Typically, a GIS software refers to data compression as a process that removes unreferenced rows from geodatabase system tables and user delta tables. Compression helps maintain versioned geodatabase performance.