Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Degree and centrality-based approaches in network-based variable selection: Insights from the Singapore Longitudinal Aging Study.

Jesus Felix Bayta Valenzuela | Christopher Monterola | Victor Joo Chuan Tong | Tamàs Fülöp | Tze Pin Ng | Anis Larbi
PloS one | 2019

We describe a network-based method to obtain a subset of representative variables from clinical data of subjects of the second Singapore Longitudinal Aging Study (SLAS-2), while preserving to a good extent the predictive performance of the full set with regards to a multi-faceted index of successful aging, SAGE. To examine differences in predictive performance of high-degree nodes ("hubs") and high-centrality ones ("cores"), we implement four subsetting strategies (two degree-based, two centrality-based) and obtain four surrogate sets of variables, which we use as input features for machine learning models to predict the SAGE index of subjects. All four models have variables belonging to the physical, cardiovascular, cognitive and immunological domains among their fifteen most important predictors. A fifth domain (leisure-time activities, LTA) is also present in some form. From a comparison of the surrogate sets' size and predictive performance, a centrality-based approach (selection of the most central variable-nodes within each cluster) yielded the smallest-sized surrogate set, while having high prediction accuracy (measured by its model's area-under-curve, AUC) in comparison to its analogous degree-based strategy (selection of the highest-degree nodes per cluster). Inclusion of the next most-central variables yielded negligible changes in predictive performance while more than doubling the surrogate set size. The centrality-based approach thus yields a surrogate set which offers a good balance between number of variables and prediction performance, and can act as a representative subset of the SLAS-2 clinical dataset.

Pubmed ID: 31318894

Research resources used in this publication

None found

Additional research tools detected in this publication

Antibodies used in this publication

None found

Associated grants

None

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


Gephi (tool)

RRID:SCR_004293

Open-source software for network visualization and analysis helping data analysts to intuitively reveal patterns and trends, highlight outliers and tells stories with their data. It uses a 3D render engine to display large graphs in real-time and to speed up the exploration. Gephi combines built-in functionalities and flexible architecture to: explore, analyze, spatialize, filter, cluterize, manipulate and export all types of networks. Gephi runs on Windows, Linux and Mac OS X. Gephi is based on a visualize-and-manipulate paradigm which allow any user to discover networks and data properties. Moreover, it is designed to follow the chain of a case study, from data file to nice printable maps. It is open-source and free (GNU General Public License). Applications: * Exploratory Data Analysis: intuition-oriented analysis by networks manipulations in real time. * Link Analysis: revealing the underlying structures of associations between objects, in particular in scale-free networks. * Social Network Analysis: easy creation of social data connectors to map community organizations and small-world networks. * Biological Network analysis: representing patterns of biological data. * Poster creation: scientific work promotion with hi-quality printable maps. Gephi 0.7 architecture is modular and therefore allows developers to add and extend functionalities with ease. New features like Metrics, Layout, Filters, Data sources and more can be easily packaged in plugins and shared. The built-in Plugins Center automatically gets the list of plugins available from the Gephi Plugin portal and takes care of all software updates. Download, comment, and rate plugins provided by community members and third-party companies, or post your own contributions!

View all literature mentions

SciPy (tool)

RRID:SCR_008058

A Python-based environment of open-source software for mathematics, science, and engineering. The core packages of SciPy include: NumPy, a base N-dimensional array package; SciPy Library, a fundamental library for scientific computing; and IPython, an enhanced interactive console.

View all literature mentions

SAGE (tool)

RRID:SCR_009302

Software application that provides researchers with the tools necessary for various types of statistical genetic analysis of human family data. (entry from Genetic Analysis Software)

View all literature mentions