NetApp Community Update
This site will enter Read Only mode on July 23 as we prepare to move to a new platform. You will still be able to view content, but posting and replying will be temporarily disabled.
We're excited to launch our new Community experience on July 30 and more information will follow soon.
Stay connected during the transition - Join our Discord community today.

Tech ONTAP Blogs

Connecting file data with AWS AI and analytics services: A governance-first approach

robertbell
NetApp
14 Views

Enterprise organizations have been storing enterprise file‑based data for decades. Amazon Web Services (AWS) has introduced a range of AI and analytics services that unlock new ways to extract insights, automate decisions, and build differentiated applications.

There is value in making on-premises file data available to these services, in terms of unlocking deeper insights. However, implementing this effectively requires a more nuanced approach. When line of business (LOB) teams ask their storage colleagues to use existing file datasets, they often aren't aware that these services use Amazon Simple Storage Service (Amazon S3) as their underlying storage infrastructure, and that the file data will need to be copied to this storage. This introduces new governance challenges, and responsibility for addressing those requirements lands squarely with the storage team, who need to turn a simple request to “use the data” into a solution that’s secure, auditable, and defensible over time. In this post, we look at a “governance‑first” approach that gives application and analytics teams a compliant path for allowing Amazon S3‑based services to consume file data without the added governance hassle, while keeping auditability and control front and center.

Read on as we cover:

 

A challenge that scales

For many organizations, the problem created by the growing number of access requests evolves gradually from a minor inconvenience into a complex operational challenge.

The first request to use a dataset with an analytics service is usually manageable. The second request might look similar. By the fifth or sixth request—which often comes from different teams, for different purposes—the job can become overwhelming.

At this point, the storage administrator becomes the traffic controller. Not because the governance-control technology is lacking, but because governance requirements increase as the same datasets are used for different clients or use cases, each with its own requirements. This is when organizations start looking for a cleaner access pattern. One that doesn’t sacrifice enterprise-grade governance and auditability in the race to be more innovative.

 

The need for a clean, auditable data access path

From the LOB team’s perspective, the requirement is clear: They need access to existing file data using AWS services that require Amazon S3-based access.

From the storage teams’ perspective, the requirement adds several additional layers of complexity as they are accountable for:

  • Clear ownership of the dataset
  • Well‑defined access guardrails
  • Consistent security enforcement
  • Ongoing audit readiness

For storage admins, success isn’t measured by how fast access is granted, but by how reliably that access can be controlled and monitored over time.

 

The traditional approach to providing Amazon S3 access to file data

Typically, storage teams support these use cases by creating an Amazon S3 data lake that can be accessed through object‑based workflows. This approach is well understood, and with the right controls in place, it can meet enterprise governance policies.

While there's nothing inherently wrong with this, a challenge emerges as more teams and use cases are added, increasing the operational burden for the storage administrators, because they must:

  • Maintain consistent access policies
  • Monitor and review usage
  • Support audit and security reviews

This ongoing supervision becomes onerous, especially as access patterns evolve.

 

A “governance‑first” architecture approach with Amazon FSx for NetApp ONTAP

This is where Amazon FSx for NetApp ONTAP (FSx for ONTAP) can help.

Because FSx for ONTAP supports multiprotocol access, the same dataset can be accessed using both file protocols and Amazon S3 APIs, ensuring the existing governance mechanisms are in place, no matter the usage. This governance-first architecture separates two responsibilities clearly:

  • FSx for ONTAP remains the storage layer for file data, preserving file access, performance, familiar file semantics, and administrative control—even if the data is used through S3 APIs.
  • Amazon S3 Access Points allows the FSx for ONTAP file data to be accessed by AWS services as if it were stored in an S3 bucket, so there’s no need to copy it to an S3 data lake. In doing so, it provides a controlled access layer for these services, making it easier to align data access with organizational governance policies.

In essence, this architecture delivers a single, controlled access mechanism that builds on the storage team’s existing security and governance controls, rather than creating multiple ad-hoc access paths.

Additionally, governance-first architecture provides ongoing visibility into how access is used over time, an important aspect in enterprise environments, where storage teams are expected to support audit and review processes long after access is granted.

 

Storage Administrators as the heroes

Storage administrators are often invisible when things go well. Data is available. Applications work. Audits pass.

This governance-first architecture approach not only simplifies their job; it also highlights the role they play in enabling innovation without sacrificing controlthey’re the ones who turn exploratory use cases into something the organization can confidently build on.

A tool available to help Storage Administrators ensure governance is NetApp Workload factory™.

 

Using NetApp Workload Factory to create and manage your S3 access points

NetApp Workload Factory is a free orchestration tool designed to help deploy and manage wellarchitected FSx for ONTAP storage.

With Workload Factory, you can both:

  • Manage your S3 access points, as shown in the following screenshot:

Picture1.png

and

  • Create and attach new S3 access points, as shown in the following screenshot:

Picture2.png

Read our post about how to use NetApp Workload Factory to set up Amazon S3 Access Points for FSx for ONTAP.

Workload Factory also features a journaling capability for FSx for ONTAP S3 access points. This provides storage administrators with records that can be referenced during security reviews or audits, without changing the underlying access model.

To learn how to set up journal table infrastructure in Workload Factory, read our post on the subject.

 

Summary and next steps

As organizations seek to unlock more value from existing file datasets, the challenge shifts from whether access is possible to how access can be provided in a way that maintains governance and auditability.

A governance‑first approach using FSx for ONTAP with S3 Access Points supports innovation by storage teams, while keeping security and compliance teams comfortable.

We invite you to learn more with these resources, which include real-world scenarios and demonstrations:

Public